<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Abonia Sojasingarayar]]></title><description><![CDATA[AI and ML Insights: AI Magazine]]></description><link>https://aboniasojasingarayar.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Vgns!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1e4c4d9c-44f8-42a5-a297-f0dcea7f031b_500x500.png</url><title>Abonia Sojasingarayar</title><link>https://aboniasojasingarayar.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 18 Aug 2026 13:19:07 GMT</lastBuildDate><atom:link href="https://aboniasojasingarayar.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Abonia Sojasingarayar]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[aboniasojasingarayar@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[aboniasojasingarayar@substack.com]]></itunes:email><itunes:name><![CDATA[Abonia Sojasingarayar]]></itunes:name></itunes:owner><itunes:author><![CDATA[Abonia Sojasingarayar]]></itunes:author><googleplay:owner><![CDATA[aboniasojasingarayar@substack.com]]></googleplay:owner><googleplay:email><![CDATA[aboniasojasingarayar@substack.com]]></googleplay:email><googleplay:author><![CDATA[Abonia Sojasingarayar]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[🤗 Hugging Face Cheatsheet]]></title><description><![CDATA[5-Page Hugging Face Ecosystem Cheatsheet]]></description><link>https://aboniasojasingarayar.substack.com/p/hugging-face-cheatsheet</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/hugging-face-cheatsheet</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Mon, 27 Jul 2026 06:19:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fgA6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I recently curated this 5-page Hugging Face Ecosystem Cheatsheet &#129303;</p><p><strong>Download the cheatsheet:</strong> https://github.com/Abonia1/HuggingFace-CheatSheet</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fgA6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fgA6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png 424w, https://substackcdn.com/image/fetch/$s_!fgA6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png 848w, https://substackcdn.com/image/fetch/$s_!fgA6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png 1272w, https://substackcdn.com/image/fetch/$s_!fgA6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fgA6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png" width="836" height="1428" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1428,&quot;width&quot;:836,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:734253,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aboniasojasingarayar.substack.com/i/208081645?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fgA6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png 424w, https://substackcdn.com/image/fetch/$s_!fgA6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png 848w, https://substackcdn.com/image/fetch/$s_!fgA6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png 1272w, https://substackcdn.com/image/fetch/$s_!fgA6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d4a786-c12e-4b72-ade3-893f86b9a31b_836x1428.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why I Created This Cheatsheet</h2><p>While working with pretrained models, datasets, fine-tuning workflows, inference systems, and deployment tools, I found that the Hugging Face ecosystem covers many different parts of the machine learning lifecycle.</p><p>There are libraries for specific tasks, platforms for sharing and versioning models and datasets, tools for training and fine-tuning, inference runtimes, deployment options, and frameworks for generative models and agents.</p><p>For someone learning the ecosystem, it can be useful to first have a high-level map of how these components relate to each other.</p><p>This is the reason I created this cheatsheet.</p><p>It is not intended to replace the official documentation or provide a complete reference to every library and API.</p><p>Instead, I tried to curate the core concepts into five pages that can be used as a visual reference.</p><div><hr></div><h2>What the Cheatsheet Covers</h2><h3>1. Hugging Face Ecosystem &amp; Hub Architecture</h3><p>The first page introduces the main components of the Hugging Face ecosystem:</p><p><strong>Hub &#8594; Models &#8594; Datasets &#8594; Spaces &#8594; Organizations &#8594; Collections &#8594; Papers &#8594; Inference &#8594; Leaderboards</strong></p><p>It also covers model repository anatomy, common model artifacts, revisions, Dataset Cards, model discovery, filtering, licensing, hardware compatibility, and authentication.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Wxob!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Wxob!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!Wxob!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!Wxob!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!Wxob!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Wxob!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png" width="1456" height="1124" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1124,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:483698,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://aboniasojasingarayar.substack.com/i/208081645?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Wxob!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!Wxob!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!Wxob!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!Wxob!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0c16f409-7bf3-4bea-91a6-0dea02619a73_1466x1132.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><h3>2. Transformers &amp; Model Execution</h3><p>The second page focuses on the execution path of pretrained transformer architectures.</p><p>The basic flow is:</p><p><strong>Input &#8594; Tokenizer / Processor &#8594; Model &#8594; Output</strong></p><p>It covers Auto Classes, task-specific model heads, tokenization, padding, truncation, sequence-length constraints, multimodal processors, Pipelines, direct model APIs, generation configuration, and chat templates.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G_o_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G_o_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!G_o_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!G_o_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!G_o_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G_o_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png" width="1456" height="1124" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1124,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:468624,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://aboniasojasingarayar.substack.com/i/208081645?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!G_o_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!G_o_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!G_o_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!G_o_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1d6313e-3b10-4575-b29a-058847328b20_1466x1132.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><h3>3. Datasets &amp; Training Stack</h3><p>The third page focuses on the data and training workflow.</p><p>It covers dataset loading, dataset splits, filtering, mapping, preprocessing, streaming, tokenization, batching, Data Collators, <code>Trainer</code>, <code>TrainingArguments</code>, <code>Accelerate</code>, PEFT, LoRA, QLoRA, evaluation, and Hub integration.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iT0g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iT0g!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!iT0g!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!iT0g!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!iT0g!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iT0g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png" width="1456" height="1124" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1124,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:469097,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://aboniasojasingarayar.substack.com/i/208081645?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!iT0g!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!iT0g!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!iT0g!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!iT0g!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0f8345-f0c4-4765-a8c9-d64b07e47001_1466x1132.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h3>4.  Inference Stack &amp; Deployment</h3><p>The fourth page focuses on how model artifacts can be executed and served.</p><p>It covers local inference, <code>InferenceClient</code>, Text Generation Inference, continuous batching, streaming, tensor parallelism, Dedicated Endpoints, quantization, BitsAndBytes, AWQ, GPTQ, Optimum, hardware optimization, vLLM, llama.cpp, and Spaces.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wGnY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wGnY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!wGnY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!wGnY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!wGnY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wGnY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png" width="1456" height="1124" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1124,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:446345,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://aboniasojasingarayar.substack.com/i/208081645?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wGnY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!wGnY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!wGnY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!wGnY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa773fc71-d16b-4d36-af3e-6b5ee05953ea_1466x1132.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h3>5. Agents, Generative Models &amp; Hub Tooling</h3><p>The final page covers several higher-level components of the ecosystem.</p><p>It includes <code>smolagents</code>, Diffusers, TRL, post-training and alignment methods, multimodal architectures, Hub CLI workflows, and Python utilities such as <code>hf_hub_download</code> and <code>snapshot_download</code>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Sfut!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Sfut!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!Sfut!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!Sfut!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!Sfut!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Sfut!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png" width="1456" height="1124" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1124,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:598646,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://aboniasojasingarayar.substack.com/i/208081645?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Sfut!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png 424w, https://substackcdn.com/image/fetch/$s_!Sfut!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png 848w, https://substackcdn.com/image/fetch/$s_!Sfut!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png 1272w, https://substackcdn.com/image/fetch/$s_!Sfut!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe133a3c5-dde5-44be-ab0d-801197e2ef51_1466x1132.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>The Structure</h2><p>The five pages are organized around a simplified development lifecycle:</p><p><strong>Discover &#8594; Load &#8594; Prepare &#8594; Train &#8594; Fine-Tune &#8594; Evaluate &#8594; Optimize &#8594; Serve &#8594; Build</strong></p><p>The purpose of this structure is to provide a starting point for understanding where different Hugging Face tools and services fit within a broader workflow.</p><div><hr></div><h2>Who Might Find It Useful?</h2><p>I hope this cheatsheet can be useful for:</p><ul><li><p>Machine Learning Engineers</p></li><li><p>LLM Engineers</p></li><li><p>Research Engineers</p></li><li><p>Applied Scientists</p></li><li><p>Data Scientists working with pretrained models</p></li><li><p>Students and practitioners learning the Hugging Face ecosystem</p></li></ul><p>It can be used as:</p><ul><li><p>A quick reference while working with Hugging Face tools</p></li><li><p>A learning aid for understanding the ecosystem</p></li><li><p>A teaching resource</p></li><li><p>A starting point for exploring the official documentation in more depth</p></li></ul><div><hr></div><h2>Access the Cheatsheet</h2><p>The complete cheatsheet is available here:</p><p>&#128279; <strong>Download:</strong> https://github.com/Abonia1/HuggingFace-CheatSheet</p><p>I hope this curated reference helps make the Hugging Face ecosystem a little easier to navigate and provides a useful starting point for further learning.</p><p>&#129303;</p><div><hr></div><h3><strong>Happy Learning!</strong></h3><div><hr></div><h1><strong>Connect with Me</strong></h1><p>If you have any inquiries, feel free to reach out via message or email.</p><blockquote><blockquote><p><em><strong><a href="https://abonia1.github.io/">Website/Newletter</a></strong></em></p><p><em><span>Connect with me on</span><strong><span> </span><a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p><p><em><span>Find me on</span><strong><span> </span><a href="https://github.com/Abonia1">Github</a></strong></em></p><p><em><span>Visit my technical channel on </span><strong><a href="https://www.youtube.com/@AboniaSojasingarayar">Youtube</a></strong></em></p></blockquote></blockquote><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[📚 E-Book - From Research to Production - Industrializing NLP and Large Language Models ]]></title><description><![CDATA[E-book release : From Research to Production - Industrializing NLP and Large Language Model]]></description><link>https://aboniasojasingarayar.substack.com/p/e-book-from-research-to-production</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/e-book-from-research-to-production</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Mon, 21 Apr 2025 07:30:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SXho!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>After a year and a half of dedicated work, I&#8217;m truly excited to share that the <strong>eBook </strong><em><strong>From Research to Production &#8211; Industrializing NLP and Large Language Models</strong></em> is now ready for release!</p><p>This free resource is a comprehensive guide covering everything from core NLP techniques to the latest advancements in large language models. This eBook is designed for beginners, intermediates, and professionals alike &#8212; or anyone simply looking to deepen their understanding. It&#8217;s our hope that this free resource supports and inspires you on your journey in the world of NLP and large language models.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SXho!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SXho!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png 424w, https://substackcdn.com/image/fetch/$s_!SXho!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png 848w, https://substackcdn.com/image/fetch/$s_!SXho!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png 1272w, https://substackcdn.com/image/fetch/$s_!SXho!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SXho!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png" width="512" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:512,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:87183,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aboniasojasingarayar.substack.com/i/154198450?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SXho!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png 424w, https://substackcdn.com/image/fetch/$s_!SXho!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png 848w, https://substackcdn.com/image/fetch/$s_!SXho!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png 1272w, https://substackcdn.com/image/fetch/$s_!SXho!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3658c90a-8967-44c1-82fa-11208efe0c52_512x800.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><p>To give you easy access to all the valuable insights shared in the eBook, I&#8217;ve structured it into a series of chapters that tackle core concepts, advanced techniques, real-world applications and more.</p><h2>Table of Content</h2><p><strong><a href="https://aboniasojasingarayar.substack.com/p/introduction-to-nlp?r=92g95">Chapter 1: Introduction to NLP</a></strong><br>Introduces the fundamentals of NLP, covering its definition, key applications, challenges, and resources for further learning.</p><p><strong><a href="https://open.substack.com/pub/aboniasojasingarayar/p/chapter-2-traditional-and-modern?r=92g95&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=false">Chapter 2: Traditional and Modern Text Representation Techniques</a></strong><br>Explores traditional text representation methods like Bag-of-Words and TF-IDF, as well as modern techniques like Word2Vec and BERT.</p><p><strong><a href="https://open.substack.com/pub/aboniasojasingarayar/p/chapter-3-advanced-language-modeling?r=92g95&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=false">Chapter 3 - Advanced Language Modeling and Transformers</a></strong><br>Diving into language models, sequence architectures, and transformer-based innovations</p><p><strong><a href="https://open.substack.com/pub/aboniasojasingarayar/p/introduction-to-large-language-models?r=92g95&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=false">Chapter 4: Introduction to Large Language Models</a></strong><br>Provides an overview of LLMs, their types (e.g., GPT-3, BART), key technical concepts, and their real-world applications like text generation and translation.</p><p><strong><a href="https://open.substack.com/pub/aboniasojasingarayar/p/chapter-3-fine-tuning-adaptation?r=92g95&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=false">Chapter 5: Fine-Tuning, Adaptation, Evaluation, and Debugging of LLMs</a></strong><br>Discusses techniques for fine-tuning LLMs, adapting them for specific tasks, evaluating their performance, and debugging models.</p><p><strong><a href="https://open.substack.com/pub/aboniasojasingarayar/p/chapter-4-real-world-applications?r=92g95&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=false">Chapter 6: Real-World Applications of RAG and LLMs</a></strong><br>Demonstrates how retrieval-augmented generation (RAG) and LLMs are used in various industries, including conversational AI, biomedical document understanding, and legal search.</p><p><strong><a href="https://aboniasojasingarayar.substack.com/p/chapter-7-mastering-llmops-operationalizing">Chapter 7: A Comprehensive Guide to LLMOps: Leveraging Large Language Models in the Era of Generative AI</a></strong><a href="https://aboniasojasingarayar.substack.com/p/chapter-7-mastering-llmops-operationalizing"> </a><br>Introduces the concept of LLMOps, which focuses on the operationalization of Large Language Models across their lifecycle, covering areas like data management, model optimization, deployment, monitoring, and security. </p><div><hr></div><h3><strong>What&#8217;s Inside the eBook</strong></h3><p>This eBook offers an in-depth exploration of text representation techniques, starting with traditional methods like Bag-of-Words and TF-IDF, and progressing to advanced models such as Word2Vec and BERT. It covers the transformative role of transformers and attention mechanisms, explaining how they have revolutionized Natural Language Processing (NLP). The eBook delves into the intricacies of transformer architecture and its applications in large-scale models like GPT and BERT, providing a solid foundation for understanding modern NLP breakthroughs. Additionally, it includes practical case studies, illustrating the application of NLP and large language models (LLMs) in various fields, including conversational AI, biomedical research, and legal domain. The guide also offers valuable insights into fine-tuning, evaluation, and debugging techniques for optimizing LLMs for specific tasks. Lastly, it introduces the emerging field of LLMOps, focusing on the best practices for deploying and maintaining large language models in real-world environments, ensuring readers are equipped with both theoretical knowledge and practical skills.</p><div><hr></div><h3><strong>Who This Book Is For</strong></h3><p>This eBook is designed for anyone interested in understanding and mastering the concepts surrounding Natural Language Processing (NLP) and Large Language Models (LLMs) and LLMOps. Whether you're a beginner or an experienced professional, this guide offers something for everyone.</p><h3><strong>Ideal for</strong></h3><ul><li><p><strong>Data Scientists and Machine Learning Engineers</strong>: If you're working in AI, this book will help you deepen your knowledge of NLP techniques, transformer-based models, and the latest advancements in LLMs, providing practical insights on how to leverage them in real-world applications.</p></li><li><p><strong>AI Enthusiasts and Researchers</strong>: For those passionate about AI and its capabilities, this guide will give you a comprehensive understanding of NLP from its traditional roots to cutting-edge LLM technologies, along with their emerging applications.</p></li><li><p><strong>Software Developers</strong>: This eBook will be a valuable resource if you're looking to incorporate NLP or LLMs into your software development projects, providing you with clear explanations of the key concepts, methods, and tools.</p></li><li><p><strong>Students and Learners</strong>: If you're just starting your journey into NLP and LLMs, this book is a great resource for gaining both foundational knowledge and advanced techniques in the field, with a focus on practical application and hands-on examples.</p></li><li><p><strong>Business Leaders and Decision-Makers</strong>: If you're considering integrating NLP and LLMs into your products or services, this eBook will help you understand the technologies and the potential benefits they offer, enabling you to make informed decisions about their implementation.</p></li></ul><div><hr></div><h3>Happy Reading!</h3><div><hr></div><h1><strong>Connect with Me</strong></h1><p>If you have any inquiries, feel free to reach out via message or email.</p><blockquote><p><em><strong><a href="https://abonia1.github.io/">Website/Newletter</a></strong></em></p><p><em>Connect with me on<strong> <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p><p><em>Find me on<strong> <a href="https://github.com/Abonia1">Github</a></strong></em></p><p><em>Visit my technical channel on <strong><a href="https://www.youtube.com/@AboniaSojasingarayar">Youtube</a></strong></em></p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[CHAPTER 7 - Mastering LLMOps: Operationalizing Large Language Models in the Generative AI Landscape]]></title><description><![CDATA[Streamlining Generative AI Workflows with LLMOps]]></description><link>https://aboniasojasingarayar.substack.com/p/chapter-7-mastering-llmops-operationalizing</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/chapter-7-mastering-llmops-operationalizing</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Mon, 24 Mar 2025 08:02:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9EYp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Large Language Models (LLMs), such as OpenAI&#8217;s GPT-4 and Google&#8217;s LaMDA, have ushered in a new era of generative AI. These models are reshaping industries&#8212;from customer engagement to creative content generation&#8212;by providing advanced natural language processing and generation capabilities. However, translating the power of LLMs into functional, real-world applications requires not just cutting-edge technology but also robust operational frameworks.</p><p>Enter <strong>LLMOps (Large Language Model Operations)</strong>: a specialized discipline focused on the lifecycle management of LLMs, bridging the gap between innovative AI research and enterprise-scale implementation. LLMOps encompasses everything from data preparation and model customization to deployment and continuous monitoring. It ensures that LLMs are not only deployed efficiently but also remain reliable, scalable, and aligned with business objectives.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Key components of LLMOps include:</p><ol><li><p><strong>Data Management</strong>: Creating high-quality, diverse datasets for fine-tuning and maintenance.</p></li><li><p><strong>Model Optimization</strong>: Tailoring pre-trained models to specific domains and tasks.</p></li><li><p><strong>Deployment Strategies</strong>: Leveraging containerization, serverless computing, and API integrations.</p></li><li><p><strong>Monitoring and Evaluation</strong>: Continuously tracking performance, bias, and fairness.</p></li><li><p><strong>Security and Compliance</strong>: Mitigating vulnerabilities and adhering to regulatory standards.</p></li></ol><p>With rapid advancements in generative AI, LLMOps practitioners must stay ahead by mastering trends like federated learning, multimodal capabilities, and explainable AI. This chapter explores the principles, challenges, and tools integral to LLMOps, equipping organizations with the knowledge to transform LLM potential into practical, scalable solutions that deliver measurable value.</p><div><hr></div><h3>What is LLMOps?</h3><p><strong>LLMOps</strong> refers to the tools, practices, and techniques used to operationalize LLMs effectively across their lifecycle. It is the generative AI equivalent of MLOps but adapted to address the unique demands of LLMs, such as massive datasets, iterative customization, and complex monitoring.</p><p>LLMOps brings together data scientists, ML engineers, and IT professionals, enabling them to deploy, monitor, and scale LLMs for consistent and high-quality performance</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9EYp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9EYp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png 424w, https://substackcdn.com/image/fetch/$s_!9EYp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png 848w, https://substackcdn.com/image/fetch/$s_!9EYp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png 1272w, https://substackcdn.com/image/fetch/$s_!9EYp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9EYp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png" width="1010" height="736" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:736,&quot;width&quot;:1010,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:199785,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9EYp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png 424w, https://substackcdn.com/image/fetch/$s_!9EYp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png 848w, https://substackcdn.com/image/fetch/$s_!9EYp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png 1272w, https://substackcdn.com/image/fetch/$s_!9EYp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86ca8d5-74f5-400b-b3ea-a7d312e2c6ea_1010x736.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h3>Why LLMOps is Critical Beyond MLOps</h3><p>While LLMOps builds on MLOps principles, it diverges significantly to meet the specific needs of generative AI. Traditional MLOps workflows often fall short when applied to LLMs, necessitating tailored approaches in areas like:</p><ol><li><p><strong>LLM Customization</strong></p><ul><li><p>Fine-tuning, prompt engineering, and retrieval-augmented generation (RAG) are iterative and resource-intensive processes, unique to LLM workflows.</p></li></ul></li><li><p><strong>Data Complexity</strong></p><ul><li><p>LLMs require vast, diverse datasets, which necessitate advanced pipelines for ingestion, transformation, and vector database integration.</p></li></ul></li><li><p><strong>Performance Monitoring</strong></p><ul><li><p>Metrics for LLMs extend beyond accuracy to include fairness, bias, and safety, requiring sophisticated monitoring tools.</p></li></ul></li><li><p><strong>Security Challenges</strong></p><ul><li><p>Unique vulnerabilities like prompt-based attacks and data poisoning necessitate advanced security measures.</p></li></ul></li></ol><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!I0bB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!I0bB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!I0bB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!I0bB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!I0bB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!I0bB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e4a197a2-70b1-4b91-a024-55f491528962_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:81346,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!I0bB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!I0bB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!I0bB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!I0bB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe4a197a2-70b1-4b91-a024-55f491528962_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2><strong>The LLMOps Lifecycle: A Step-by-Step Process</strong></h2><p>To operationalize LLMs successfully, LLMOps introduces a structured lifecycle encompassing <strong>exploratory data analysis, customization, deployment, and monitoring</strong>. Below is a detailed breakdown:</p><p><strong>1. Exploratory Data Analysis (EDA)</strong></p><ul><li><p><strong>Data Understanding:</strong> Analyze the characteristics of the dataset, identifying patterns, outliers, or gaps.</p></li><li><p><strong>Data Collection:</strong> Gather information from diverse sources aligned with the LLM&#8217;s intended use case.</p></li><li><p><strong>Data Cleaning:</strong> Eliminate errors, inconsistencies and duplicates to prepare a high-quality dataset.</p></li></ul><p><strong>2. Model Selection and Customization</strong></p><ul><li><p><strong>Model Selection:</strong> Choose an LLM architecture (e.g., GPT, BERT) based on the task requirements and available resources.</p></li><li><p><strong>Customization Methods:</strong></p><ul><li><p><strong>Fine-Tuning:</strong> Adjust model parameters with domain-specific data for specialized tasks.</p></li><li><p><strong>Prompt Tuning:</strong> Use optimized prompt templates to guide model outputs without altering its weights.</p></li><li><p><strong>Retrieval-Augmented Generation (RAG):</strong> Leverage external databases to enrich responses with contextual information.</p></li></ul></li></ul><p><strong>3. Model Deployment</strong></p><ul><li><p>Prepare the customized LLM for <strong>serving in a production environment</strong>, ensuring efficient infrastructure setup.</p></li><li><p>Deploy incrementally, starting with quality assurance (QA) environments, before moving to production.</p></li></ul><p><strong>4. Ongoing Monitoring</strong></p><ul><li><p>Monitor critical metrics, including latency, accuracy, cost, and fairness, to detect performance degradation/drifts.</p></li><li><p>Implement real-time alerts for issues such as data drift or output quality changes.</p></li></ul><div><hr></div><h2><strong>Tools for Implementing LLMOps</strong></h2><p>A wide array of tools is available to support LLMOps across its lifecycle. Some popular options include:</p><h3><strong>Data Engineering</strong></h3><p>Data engineering forms the backbone of any LLMOps workflow, ensuring that the model is powered by high-quality, relevant data. This involves data ingestion, transformation, and integration with advanced storage and retrieval systems.</p><h4><strong>Popular Databases for LLMOps</strong></h4><p>LLMOps requires robust databases for storing, managing, and retrieving large-scale datasets and embeddings. Popular options include:</p><ul><li><p><strong><a href="https://www.pinecone.io/">Pinecone</a></strong>: A vector database optimized for retrieval-augmented generation (RAG), allowing efficient similarity searches for embeddings. </p></li><li><p><strong><a href="https://weaviate.io/">Weaviate</a></strong>: An open-source vector database with integrated machine learning capabilities for seamless data storage and retrieval. </p></li><li><p><strong><a href="https://github.com/facebookresearch/faiss">FAISS (Facebook AI Similarity Search)</a>: </strong>A library that enables fast and scalable similarity searches across embeddings. </p></li><li><p><strong><a href="https://milvus.io/">Milvus</a></strong>: A highly scalable vector database that supports real-time embedding management and queries. </p></li><li><p><strong><a href="https://www.elastic.co/elasticsearch/">ElasticSearch</a></strong>: A versatile search engine that supports text and vector-based queries, commonly used in LLMOps for hybrid search capabilities. </p></li><li><p><strong><a href="https://www.postgresql.org/">PostgreSQL</a></strong>: A robust relational database often extended with vector search plugins like pgvector for hybrid storage needs. </p></li></ul><p><strong>Vector Search with Pinecone</strong></p><pre><code>import llama_index

from llama_index.vector_stores.pinecone import PineconeVectorStore

from llama_index.core import StorageContext, SimpleDirectoryReader

# Initialize Pinecone

pinecone.init(api_key="api_key", environment="your_environment")

# Create a new index

index_name = "example-index"

pinecone.create_index(index_name, dimension=128)

# Construct vector store

vector_store = PineconeVectorStore(pinecone_index=pinecone.Index(index_name))

# Create storage context

storage_context = StorageContext.from_defaults(vector_store=vector_store)

# Load documents

documents = SimpleDirectoryReader("../data").load_data()

# Build index

index = llama_index.VectorStoreIndex.from_documents(

documents,

storage_context=storage_context,

)

# Query the index

query_engine = index.as_query_engine()

response = query_engine.query("Your search query here")

print(response)

Vector Search with Faiss

import numpy as np

from llama_index.vector_stores.faiss import FaissVectorStore

from llama_index.core import StorageContext, SimpleDirectoryReader

# Create FAISS index

dimension = 128

faiss_index = faiss.IndexFlatL2(dimension)

# Construct vector store

faiss_vector_store = FaissVectorStore(faiss_index)

# Create storage context

storage_context = StorageContext.from_defaults(vector_store=faiss_vector_store)

# Load documents

documents = SimpleDirectoryReader("../data").load_data()

# Build index

index = llama_index.VectorStoreIndex.from_documents(

documents,

storage_context=storage_context,

)

# Perform similarity search

query_vector = np.random.random((1, dimension)).astype("float32")

D, I = index.storage_context.vector_store.search(query_vector, k=3)

print("Distances:", D)

print("Indices:", I)</code></pre><h4><strong>Popular Data Crawlers for LLMOps</strong></h4><p>Efficient data collection often involves using crawlers to gather information from diverse sources. Leading tools include:</p><ul><li><p><strong><a href="https://scrapy.org/">Scrapy</a></strong>: A Python-based web scraping framework designed for scalability and flexibility in crawling structured and unstructured data. </p></li><li><p><strong><a href="https://www.crummy.com/software/BeautifulSoup/">BeautifulSoup</a></strong>: A library for extracting data from HTML and XML files, ideal for quick, small-scale crawling tasks. </p></li><li><p><strong><a href="https://nutch.apache.org/">Apache Nutch</a></strong>: An open-source web crawler that integrates seamlessly with big data tools like Hadoop. </p></li><li><p><strong><a href="https://www.octoparse.com/">Octoparse</a></strong>: A no-code web scraping tool that simplifies data extraction for users with limited programming knowledge. </p></li><li><p><strong><a href="https://diffbot.com/">Diffbot</a></strong>: A powerful API-based data extraction tool that can scrape web pages and transform content into structured formats. </p></li><li><p><strong><a href="https://github.com/gocolly/colly">Colly</a></strong>: A fast, scalable, and elegant crawler written in Go, known for its simplicity and high performance. </p></li></ul><pre><code>import scrapy

from llama_index.core.data_loader import DataLoader

from llama_index.core.schema import Node

class CustomDataLoader(DataLoader):

def __init__(self, spider_class):

self.spider = spider_class()

def load_data(self, batch_size=100):

for i in range(0, len(self.spider.start_urls), batch_size):

urls = self.spider.start_urls[i:i+batch_size]

items = []

for url in urls:

item = self.spider.parse(url)

items.append(item)

yield items

# Define your Scrapy Spider class here

class QuotesSpider(scrapy.Spider):

name = "quotes"

def start_requests(self):

urls = ["http://quotes.toscrape.com"]

for url in urls:

yield scrapy.Request(url=url, callback=self.parse)

def parse(self, response):

for quote in response.css("div.quote"):

yield {

"text": quote.css("span.text::text").get(),

"author": quote.css("span small.author::text").get(),

}

# Use the custom data loader with LlamaIndex

custom_loader = CustomDataLoader(QuotesSpider)

documents = custom_loader.load_data()

index = llama_index.VectorStoreIndex.from_documents(documents)</code></pre><div><hr></div><h3><strong>LLM Libraries - Prompt optimization, semantic search, and retrieval</strong></h3><ul><li><p><strong><a href="https://python.langchain.com/">LangChain</a></strong>: Provides building blocks for LLM-powered applications, including tools for prompt optimization, version control, and deployment. </p></li><li><p><strong><a href="https://github.com/deepset-ai/haystack">Haystack</a></strong>: Facilitates semantic search, question-answering, and LLM agent design. </p></li><li><p><strong><a href="https://langgraph.ai/">LangGraph</a></strong>: Framework for developing AI agents using graph-based representations. </p></li><li><p><strong><a href="https://github.com/langchain/llm-langserve">LangServe</a></strong>: Library for deploying LangChain applications via REST API. </p></li><li><p><strong><a href="https://llama-index.github.io/">LlamaIndex</a></strong>: Comprehensive toolkit for building LLM applications, including vector stores and retrieval. </p></li></ul><div><hr></div><h3><strong>Monitoring and Observability</strong></h3><ul><li><p><strong><a href="https://langsmith.org/">LangSmith</a></strong>: Offers robust observability features, including token counting and performance monitoring, to optimize LLM deployments. </p></li><li><p><strong><a href="https://github.com/langchain/llm-langfuse">Langfuse</a></strong>: Open-source observability platform for LLM applications, providing detailed traces and evaluation metrics. </p></li><li><p><strong><a href="https://www.tensorflow.org/tfx">TensorFlow Extended (TFX)</a></strong>: Machine learning lifecycle management platform. </p></li><li><p><strong><a href="https://mlflow.org/">MLflow</a></strong>: Open-source platform for the machine learning lifecycle, including experiment tracking and model versioning. </p></li><li><p><strong><a href="https://greatexpectations.io/">Great Expectations</a></strong>: Library for data validation and testing. </p><div><hr></div></li></ul><h3><strong>Data Management and Security</strong></h3><ul><li><p><strong><a href="https://dvc.org/">DVC (Data Version Control)</a></strong>: Tool for version controlling machine learning projects. </p></li><li><p><strong><a href="https://delta.io/">Delta Lake:</a></strong> Framework for managing data lakes. </p></li><li><p><strong><a href="https://airflow.apache.org/">Apache Airflow</a>:</strong> Platform for orchestrating complex computational workflows. </p></li><li><p><strong><a href="https://www.hashicorp.com/products/vault">HashiCorp</a></strong> Vault: Secrets management solution. </p></li><li><p><strong>AWS</strong> Key Management Service (<strong>KMS</strong>) / <strong>Azure</strong> Key Vault / <strong>Google</strong> Cloud KMS: Encryption key management services.</p><div><hr></div></li></ul><h3><strong>Model Selection and Benchmarking</strong></h3><ul><li><p><strong><a href="https://huggingface.co/models">Hugging Face Model Hub</a></strong>: Repository of pre-trained models and benchmarking tools. </p></li><li><p><strong><a href="https://llama-index.github.io/">LlamaIndex</a></strong>: Framework for benchmarking and comparing models across various tasks. </p></li><li><p><strong><a href="https://mlperf.org/">MLPerf</a></strong>: Industry standard benchmark suite for AI systems. </p></li><li><p><strong><a href="https://www.deepchecks.org/">Deepchecks</a></strong>: Comprehensive model evaluation platform. </p></li></ul><div><hr></div><h3><strong>Version Control and Deployment</strong></h3><ul><li><p><strong><a href="https://git-scm.com/">Git</a></strong>: Distributed version control system. </p></li><li><p><strong><a href="https://www.docker.com/">Docker</a></strong>: Containerization platform for deploying applications.</p></li><li><p><strong><a href="https://kubernetes.io/">Kubernetes</a></strong>: Container orchestration system. </p></li><li><p><strong><a href="https://flask.palletsprojects.com/">Flask</a></strong> / <strong><a href="https://www.djangoproject.com/">Django</a></strong> / <strong><a href="https://fastapi.tiangolo.com/">FastAPI</a></strong>: Frameworks for building RESTful APIs. </p></li></ul><div><hr></div><h3>Best Practices for Effective LLMOps</h3><ol><li><p><strong>Data Management and Security</strong></p><ul><li><p>Adopt rigorous pipelines for data preprocessing and encryption.</p></li><li><p>Use tools like <strong>Apache Airflow</strong>, <strong>DVC</strong>, and <strong>Delta Lake</strong>.</p></li></ul></li><li><p><strong>Model Customization and Selection</strong></p><ul><li><p>Choose models based on task-specific needs using benchmarking tools like <strong>Hugging Face</strong> and <strong>LlamaIndex</strong>.</p></li></ul></li><li><p><strong>Scalable Deployment</strong></p><ul><li><p>Implement containerized solutions with tools like <strong>Kubernetes</strong>.</p></li><li><p>Opt for serverless deployments using <strong>AWS Lambda</strong> or <strong>Google Cloud Functions</strong>.</p></li></ul></li><li><p><strong>Continuous Monitoring</strong></p><ul><li><p>Employ real-time observability platforms like <strong>LangSmith</strong> and <strong>Langfuse</strong>.</p></li><li><p>Regularly retrain models to maintain relevance </p></li></ul><p></p></li></ol><div><hr></div><h3>The Future of LLMOps</h3><p>Key trends shaping the future of LLMOps include:</p><ol><li><p><strong>Responsible AI</strong>: Enhanced focus on ethical practices and transparency.</p></li><li><p><strong>Hybrid Approaches</strong>: Combining LLMs with other AI paradigms.</p></li><li><p><strong>Democratization of AI</strong>: Expanding access to LLM capabilities.</p></li><li><p><strong>Sustainability</strong>: Addressing the environmental impact of large-scale AI systems.</p></li></ol><p>By embracing these trends, LLMOps practitioners can turn the transformative potential of generative AI into tangible, responsible business outcomes. LLMOps isn&#8217;t just about operationalizing technology&#8212;it&#8217;s about shaping the future of AI deployment.</p><div><hr></div><h1><strong>Connect with Me</strong></h1><p>If you have any inquiries, feel free to reach out via message or email.</p><blockquote><p><em><strong><a href="https://abonia1.github.io/">Website/Newletter</a></strong></em></p><p><em>Connect with me on<strong> <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p><p><em>Find me on<strong> <a href="https://github.com/Abonia1">Github</a></strong></em></p><p><em>Visit my technical channel on <strong><a href="https://www.youtube.com/@AboniaSojasingarayar">Youtube</a></strong></em></p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Chapter 3 - Advanced Language Modeling and Transformers]]></title><description><![CDATA[Diving into language models, sequence architectures, and transformer-based innovations]]></description><link>https://aboniasojasingarayar.substack.com/p/chapter-3-advanced-language-modeling</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/chapter-3-advanced-language-modeling</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Mon, 24 Feb 2025 08:22:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!MZBV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4><strong>Table of Content</strong></h4><p><strong>3.1 Language Modeling: N-gram models, RNNs, LSTMs, GRUs</strong></p><ul><li><p><strong>3.1.1</strong> Foundations of statistical language modeling</p></li><li><p><strong>3.1.2</strong> Architectures of RNNs, LSTMs, GRUs</p></li><li><p><strong>3.3.3</strong> Training and optimization of language models</p></li></ul><p><strong>3.2 Transformers and Attention Mechanisms</strong></p><ul><li><p><strong>3.2.1</strong> Multi-headed self-attention</p></li><li><p><strong>3.2.2</strong> Transformer encoder-decoder architecture</p><p></p><div><hr></div></li></ul><p>Written in Collaboration - Special thanks to <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Manan Thakkar&quot;,&quot;id&quot;:300408543,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10e24dac-fd32-4234-ac63-24a90d0f362c_144x144.png&quot;,&quot;uuid&quot;:&quot;72f4252c-7118-4257-b530-96423d368721&quot;}" data-component-name="MentionToDOM"></span> for his valuable contributions to structure this chapter, shaping it into an informative and accessible resource for understanding Advanced Language Modeling and Transformers.</p><p>In this chapter, we  focuses on language modeling, which aims to predict likely next words given previous text. We first cover the foundations of statistical n-gram language models. Then, key neural network architectures for language modeling are presented, including recurrent neural networks, long short-term memory networks, and gated recurrent units. Special attention is given to transformers, state-of-the-art models that employ multi-headed self-attention to capture long-range dependencies in text. Pre-training strategies for transformers like BERT are also discussed.</p><p>By the end of this chapter, you will have strong conceptual and mathematical foundations regarding language modeling, enabling you to develop predictive NLP systems. The techniques presented form the basis for contemporary NLP with deep neural networks.</p><ul><li><p>Language Modeling Foundations: Statistical n-gram models and neural network architectures like RNNs, LSTMs, and GRUs</p></li><li><p>Transformers and Attention: Multi-headed self-attention mechanisms, encoder-decoder architectures, and transformer pre-training strategies</p></li><li><p>Comparative Analysis: Contrasting different techniques for text representation and language modeling in NLP</p></li><li><p>Resources: Overview of key datasets, libraries, and learning materials on text representation and language modeling</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div></li></ul><h2>Language Modeling: N-gram models, RNNs, LSTMs, GRUs</h2><p>Language modeling is a fundamental task in NLP that aims to predict the probability of a sequence of words. It plays a crucial role in various applications, such as speech recognition, machine translation, text generation, and sentiment analysis. Over the years, several techniques have been developed to tackle the challenge of language modeling, ranging from traditional statistical approaches to more advanced deep learning methods. In this section, we will explore the evolution of language modeling techniques, starting with the classic N-gram models and then delving into the powerful neural network architectures, including Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and Gated Recurrent Units (GRUs). We will discuss the strengths and limitations of each approach and understand how they have contributed to the advancement of NLP.</p><div><hr></div><h3>N-gram Models</h3><p>N-gram models are a classic statistical approach to language modeling that have been widely used in various NLP tasks. They are based on the assumption that the probability of a word depends only on the previous N-1 words, where N is a fixed number. In this section, we will explore the concept of N-gram models, their algorithmic implementation, and their advantages and limitations.</p><p>An N-gram is a contiguous sequence of N words from a given text or speech corpus. The key idea behind N-gram models is to estimate the probability of a word based on the context of the preceding N-1 words. The value of N determines the size of the context window and the order of the N-gram model. For example:</p><ul><li><p>Unigram (N=1): Considers each word independently, without any context.</p></li><li><p>Bigram (N=2): Considers the probability of a word given the previous word.</p></li><li><p>Trigram (N=3): Considers the probability of a word given the previous two words.</p></li></ul><p>The probability of a sequence of words is calculated as the product of the individual N-gram probabilities. For example, in a bigram model, the probability of a sentence "The cat sat on the mat" would be calculated as:</p><blockquote><p><code>P("The cat sat on the mat") =</code> <code>P("The") &#215; P("cat" | "The") &#215; P("sat" | "cat") &#215; P("on" | "sat") &#215; P("the" | "on") &#215; P("mat" | "the")</code></p></blockquote><p><strong>Algorithmic Implementation</strong></p><p>The process of building an N-gram model involves the following steps:</p><ul><li><p>Tokenization: Split the text corpus into individual words or tokens.</p></li><li><p>N-gram Generation: Create N-gram sequences by sliding a window of size N over the tokenized text.</p></li><li><p>Frequency Counting: Count the frequency of each N-gram in the corpus.</p></li><li><p>Probability Estimation: Calculate the probability of each N-gram by dividing its frequency by the frequency of the preceding (N-1)-gram.</p></li></ul><p><strong>Example:</strong> Predicting the Next Word using N-gram Model</p><p>Generated Corpus:</p><p><em>&lt;S&gt; NLP is fascinating &lt;/S&gt;</em></p><p><em>&lt;S&gt; ML is a subset of AI &lt;/S&gt;</em></p><p><em>&lt;S&gt; NLP and ML are important &lt;/S&gt;</em></p><p>Question: Given the bigram "&lt;S&gt; NLP", predict the most probable next word using the bigram model.</p><p><strong>Step 1: Generate the frequency table for the corpus.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MZBV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MZBV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png 424w, https://substackcdn.com/image/fetch/$s_!MZBV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png 848w, https://substackcdn.com/image/fetch/$s_!MZBV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png 1272w, https://substackcdn.com/image/fetch/$s_!MZBV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MZBV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png" width="696" height="1011" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1011,&quot;width&quot;:696,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:61938,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!MZBV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png 424w, https://substackcdn.com/image/fetch/$s_!MZBV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png 848w, https://substackcdn.com/image/fetch/$s_!MZBV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png 1272w, https://substackcdn.com/image/fetch/$s_!MZBV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7981b1f-9a91-429b-965b-eea157b8fbfe_696x1011.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Step 2: Generate the probability table for the next word given the bigram "&lt;S&gt; NLP".</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s_jL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s_jL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png 424w, https://substackcdn.com/image/fetch/$s_!s_jL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png 848w, https://substackcdn.com/image/fetch/$s_!s_jL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png 1272w, https://substackcdn.com/image/fetch/$s_!s_jL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s_jL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png" width="1324" height="1034" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1034,&quot;width&quot;:1324,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:149454,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!s_jL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png 424w, https://substackcdn.com/image/fetch/$s_!s_jL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png 848w, https://substackcdn.com/image/fetch/$s_!s_jL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png 1272w, https://substackcdn.com/image/fetch/$s_!s_jL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F84309241-2ec5-464c-ad2b-40ee037c3900_1324x1034.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Step 3: Identify the most probable next word.</strong></p><p>Based on the probability table, the most probable next words after "&lt;S&gt; NLP" are "is" and "and", both with a probability of 0.5.</p><p>Therefore, the most probable next words after "&lt;S&gt; NLP" are "is" and "and".</p><p><strong>Code:</strong></p><pre><code><code>from collections import defaultdict

import re

def train_bigram_model(text):

# Tokenize the text

words = re.findall(r'\b\w+\b', text.lower())

# Create bigrams

bigrams = [(words[i], words[i+1]) for i in range(len(words)-1)]

# Count bigram frequencies

bigram_counts = defaultdict(int)

for bigram in bigrams:

bigram_counts[bigram] += 1

# Count word frequencies

word_counts = defaultdict(int)

for word in words:

word_counts[word] += 1

# Calculate bigram probabilities

bigram_probs = {}

for bigram, count in bigram_counts.items():

word1, word2 = bigram

bigram_probs[bigram] = count / word_counts[word1]

return bigram_probs

# Example usage

text = "The cat sat on the mat. The dog chased the cat."

bigram_model = train_bigram_model(text)

# Print bigram probabilities

for bigram, prob in bigram_model.items():

print(f"{bigram}: {prob}")</code></code></pre><div><hr></div><h3>RNN</h3><p>Feedforward and convolutional neural networks (CNNs) have been successful in various tasks, but they have limitations when it comes to processing sequential data. In CNNs, the input size is fixed, and each input is treated independently of the others. For example, when feeding images to a CNN for classification, the computations and decisions for two successive images are completely independent of each other.</p><p>However, many real-world problems involve sequential data, where the inputs are not of fixed size, and successive inputs may be dependent on each other. Some examples of sequence learning problems include:</p><ul><li><p><strong>Auto-completion:</strong> Given the first character 'd', predicting the next character 'e' and so on.</p></li><li><p><strong>Part-of-speech tagging:</strong> Predicting the part of speech tag (noun, adverb, adjective, verb) of each word in a sentence.</p></li><li><p><strong>Sentiment analysis:</strong> Predicting the polarity of a movie review based on the entire sequence of words.</p></li></ul><p>In these scenarios, the current output may depend on the current input as well as the previous inputs, and the size of the input is not fixed. Traditional neural networks struggle to handle such tasks effectively.</p><p><strong>The Need for RNNs</strong></p><p>To address the limitations of feedforward and CNNs in handling sequential data, RNNs were introduced. RNNs are designed to handle tasks involving sequences by maintaining an internal state that allows them to capture and exploit dependencies between inputs.</p><p>RNNs have three main targets:</p><ul><li><p>Account for dependence between inputs</p></li><li><p>Account for variable number of inputs</p></li><li><p>Make sure that the function executed at each time step is the same</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZYlf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZYlf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png 424w, https://substackcdn.com/image/fetch/$s_!ZYlf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png 848w, https://substackcdn.com/image/fetch/$s_!ZYlf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png 1272w, https://substackcdn.com/image/fetch/$s_!ZYlf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZYlf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png" width="416" height="325" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:325,&quot;width&quot;:416,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!ZYlf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png 424w, https://substackcdn.com/image/fetch/$s_!ZYlf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png 848w, https://substackcdn.com/image/fetch/$s_!ZYlf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png 1272w, https://substackcdn.com/image/fetch/$s_!ZYlf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1539bcf9-dab4-49f0-8ac2-637fcb17c019_416x325.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure  &#8211; Basic RNN architecture</p><p><strong>Recurrent Connections in RNNs</strong></p><p>The key component of an RNN is the recurrent connection, which allows the network to maintain a hidden state that captures information from previous time steps. At each time step, the RNN takes an input and the previous hidden state, and produces an output and an updated hidden state.</p><p>The computation in an RNN can be summarized using the following equations:</p><p>where:</p><ul><li><p>s<sub>i</sub> is the hidden state at time step i</p></li><li><p>x<sub>i</sub> is the input at time step i</p></li><li><p>U, W, and V are weight matrices</p></li><li><p>b and c are bias vectors</p></li><li><p>is an activation function (e.g., sigmoid or tanh)</p></li><li><p>O is the output function (e.g., softmax for classification)</p></li></ul><p>The recurrent connection (Ws<sub>i-1</sub>) allows the hidden state to capture information from previous time steps, enabling the RNN to handle dependencies between inputs. This recurrent connection is crucial for modeling sequential data effectively.</p><p><strong>Training RNNs with Backpropagation Through Time (BPTT)</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!I-dx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!I-dx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png 424w, https://substackcdn.com/image/fetch/$s_!I-dx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png 848w, https://substackcdn.com/image/fetch/$s_!I-dx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png 1272w, https://substackcdn.com/image/fetch/$s_!I-dx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!I-dx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png" width="426" height="417" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d2881911-bf50-4c28-814d-85ce42af61f6_426x417.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:417,&quot;width&quot;:426,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!I-dx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png 424w, https://substackcdn.com/image/fetch/$s_!I-dx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png 848w, https://substackcdn.com/image/fetch/$s_!I-dx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png 1272w, https://substackcdn.com/image/fetch/$s_!I-dx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2881911-bf50-4c28-814d-85ce42af61f6_426x417.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure  &#8211; RNN with backpropagation</p><p>Training an RNN involves using the backpropagation through time (BPTT) algorithm, which unrolls the network over the sequence and applies the standard backpropagation algorithm to compute the gradients.</p><p>The total loss in an RNN is the sum of the losses over all time steps. For each time step t, the loss is computed based on the predicted output y<sub>t</sub> and the actual target c:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3piA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3piA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png 424w, https://substackcdn.com/image/fetch/$s_!3piA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png 848w, https://substackcdn.com/image/fetch/$s_!3piA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png 1272w, https://substackcdn.com/image/fetch/$s_!3piA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3piA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png" width="540" height="48" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:48,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:12942,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!3piA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png 424w, https://substackcdn.com/image/fetch/$s_!3piA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png 848w, https://substackcdn.com/image/fetch/$s_!3piA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png 1272w, https://substackcdn.com/image/fetch/$s_!3piA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15f8083d-83c2-40dd-9bfb-192819973b44_540x48.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>To update the parameters of the RNN (U, V, W), the gradients of the loss with respect to these parameters are computed using the chain rule of differentiation. The gradients are then used to update the parameters using gradient descent or its variants.</p><p><strong>Vanishing and Exploding Gradients Problem</strong></p><p>One of the challenges in training RNNs is the vanishing and exploding gradients problem. As the backpropagation algorithm advances backward from the output layer towards the input layer, the gradients can become very small (vanishing) or very large (exploding).</p><ul><li><p>Vanishing gradients occur when the gradients approach zero, leaving the weights of the initial or lower layers nearly unchanged. This prevents the network from learning long-term dependencies effectively.</p></li><li><p>Exploding gradients occur when the gradients keep getting larger, causing very large weight updates and making the gradient descent diverge.</p></li></ul><p>To mitigate these issues, techniques such as gradient clipping, using activation functions like ReLU, and more advanced architectures like LSTM or GRU can be employed.</p><p>Despite these challenges, RNNs have been widely used and have achieved significant success in various sequence modeling tasks, including language modeling, speech recognition, and machine translation. They have paved the way for more advanced architectures that further enhance the ability to capture and utilize long-term dependencies in sequential data.</p><div><hr></div><h3>LSTM</h3><p>RNNs have shown great success in handling sequential data, but they suffer from the problem of vanishing or exploding gradients when dealing with long-term dependencies. This limitation hinders the ability of RNNs to capture and utilize information from distant time steps effectively. To address this issue, LSTM networks were introduced.</p><p>LSTM is a type of recurrent neural network architecture specifically designed to handle long-term dependencies in sequential data. It introduces a memory cell and three types of gates (forget gate, input gate, and output gate) that regulate the flow of information within the network. These components enable LSTM to selectively remember or forget information over long sequences, making it capable of capturing long-term dependencies.</p><p><strong>Architectural Components</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cDzb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cDzb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png 424w, https://substackcdn.com/image/fetch/$s_!cDzb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png 848w, https://substackcdn.com/image/fetch/$s_!cDzb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png 1272w, https://substackcdn.com/image/fetch/$s_!cDzb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cDzb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png" width="575" height="300" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:300,&quot;width&quot;:575,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!cDzb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png 424w, https://substackcdn.com/image/fetch/$s_!cDzb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png 848w, https://substackcdn.com/image/fetch/$s_!cDzb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png 1272w, https://substackcdn.com/image/fetch/$s_!cDzb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfbb7489-b881-4bcd-a603-cf03c4a79218_575x300.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure  &#8211; LSTM architecture</p><p>The LSTM architecture consists of the following key components:</p><p><strong>Memory Cell (C):</strong></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Eb_7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Eb_7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png 424w, https://substackcdn.com/image/fetch/$s_!Eb_7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png 848w, https://substackcdn.com/image/fetch/$s_!Eb_7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png 1272w, https://substackcdn.com/image/fetch/$s_!Eb_7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Eb_7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png" width="540" height="72" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:72,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:11894,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!Eb_7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png 424w, https://substackcdn.com/image/fetch/$s_!Eb_7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png 848w, https://substackcdn.com/image/fetch/$s_!Eb_7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png 1272w, https://substackcdn.com/image/fetch/$s_!Eb_7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b4b59fd-3b22-4e36-a1cb-62809e9fcbbe_540x72.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><ul><li><p>The memory cell is the core component of LSTM that stores the long-term memory of the network.</p></li><li><p>It is responsible for maintaining the state information over time.</p></li><li><p>The memory cell is controlled by the forget gate, input gate, and output gate, which regulate the flow of information into and out of the cell.</p></li><li><p>The state of the memory cell is updated at each time step based on the outputs of the gates.</p></li><li><p>The memory cell update equation is given by:</p></li></ul><ul><li><p>Here, f<sub>t</sub> is the output of the forget gate, C<sub>t-1</sub> is the previous memory cell state, i<sub>t</sub> is the output of the input gate, and is the candidate memory cell.</p><p><strong>Hidden State (h):</strong></p></li></ul><ul><li><p>The hidden state represents the output of the LSTM unit at each time step.</p></li><li><p>It captures the relevant information from the current input and the previous hidden state.</p></li><li><p>The hidden state is used as input to the next time step and can also be used for making predictions or as input to other layers in the network.</p></li><li><p>The hidden state equation is given by:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KLgT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KLgT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png 424w, https://substackcdn.com/image/fetch/$s_!KLgT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png 848w, https://substackcdn.com/image/fetch/$s_!KLgT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png 1272w, https://substackcdn.com/image/fetch/$s_!KLgT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KLgT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png" width="540" height="68" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:68,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9762,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!KLgT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png 424w, https://substackcdn.com/image/fetch/$s_!KLgT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png 848w, https://substackcdn.com/image/fetch/$s_!KLgT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png 1272w, https://substackcdn.com/image/fetch/$s_!KLgT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0f3f540-3ea8-4dfb-9a9b-5a62d32ef990_540x68.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><ul><li><p>Here, o<sub>t</sub> is the output of the output gate, and C<sub>t</sub> is the updated memory cell state.</p><p><strong>Forget Gate (f):</strong></p></li></ul><ul><li><p>The forget gate determines what information to discard from the memory cell.</p></li><li><p>It takes the previous hidden state h<sub>t-1</sub> and the current input x<sub>t</sub> as inputs and produces a value between 0 and 1 for each element in the memory cell.</p></li><li><p>A value of 0 means completely forget the corresponding element, while a value of 1 means completely retain it.</p></li><li><p>The forget gate helps the LSTM to selectively forget irrelevant information from the past.</p></li><li><p>The forget gate equation is given by:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MIvm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MIvm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png 424w, https://substackcdn.com/image/fetch/$s_!MIvm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png 848w, https://substackcdn.com/image/fetch/$s_!MIvm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png 1272w, https://substackcdn.com/image/fetch/$s_!MIvm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MIvm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png" width="540" height="67" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:67,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9715,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!MIvm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png 424w, https://substackcdn.com/image/fetch/$s_!MIvm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png 848w, https://substackcdn.com/image/fetch/$s_!MIvm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png 1272w, https://substackcdn.com/image/fetch/$s_!MIvm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd23dd2-14d0-4a40-a8ec-4923f5bc5c40_540x67.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><ul><li><p>Here, W<sub>f</sub> is the weight matrix for the forget gate, h<sub>t-1</sub> is the previous hidden state, x<sub>t</sub> is the current input, b<sub>f</sub> is the bias term, and is the sigmoid activation function.</p><p><strong>Input Gate (i) and Candidate Memory Cell (</strong>C<strong>):</strong></p></li></ul><ul><li><p>The input gate decides what new information to store in the memory cell.</p></li><li><p>It combines the previous hidden state (h<sub>t-1</sub>) and the current input (x<sub>t</sub>) to produce a value between 0 and 1 for each element in the candidate memory cell (C).</p></li><li><p>The candidate memory cell represents the new information that could potentially be added to the memory cell.</p></li><li><p>The input gate controls which elements of the candidate memory cell will be updated in the memory cell.</p></li><li><p>The input gate equation is given by:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-AoB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-AoB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png 424w, https://substackcdn.com/image/fetch/$s_!-AoB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png 848w, https://substackcdn.com/image/fetch/$s_!-AoB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png 1272w, https://substackcdn.com/image/fetch/$s_!-AoB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-AoB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png" width="540" height="65" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:65,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:9492,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!-AoB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png 424w, https://substackcdn.com/image/fetch/$s_!-AoB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png 848w, https://substackcdn.com/image/fetch/$s_!-AoB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png 1272w, https://substackcdn.com/image/fetch/$s_!-AoB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a4cd741-f11d-41d7-985c-255b88509b72_540x65.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><ul><li><p>The candidate memory cell equation is given by:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s63l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s63l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png 424w, https://substackcdn.com/image/fetch/$s_!s63l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png 848w, https://substackcdn.com/image/fetch/$s_!s63l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png 1272w, https://substackcdn.com/image/fetch/$s_!s63l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s63l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png" width="540" height="72" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4418d3a4-91d2-4f52-a42c-098184833121_540x72.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:72,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:17037,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!s63l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png 424w, https://substackcdn.com/image/fetch/$s_!s63l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png 848w, https://substackcdn.com/image/fetch/$s_!s63l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png 1272w, https://substackcdn.com/image/fetch/$s_!s63l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4418d3a4-91d2-4f52-a42c-098184833121_540x72.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Here, W<sub>i</sub> and W<sub>C</sub> are the weight matrices for the input gate and candidate memory cell, respectively, and b<sub>i</sub> and b<sub>C</sub> are the corresponding bias terms.</p><p><strong>Output Gate (o):</strong></p><ul><li><p>The output gate controls what information from the memory cell to output.</p></li><li><p>It takes the previous hidden state (h<sub>t-1</sub>) and the current input (x<sub>t</sub>) as inputs and produces a value between 0 and 1 for each element in the memory cell.</p></li><li><p>The output gate determines which parts of the memory cell will be exposed to the next time step and influence the computation of the hidden state.</p></li><li><p>The output gate equation is given by:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tpKi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tpKi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png 424w, https://substackcdn.com/image/fetch/$s_!tpKi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png 848w, https://substackcdn.com/image/fetch/$s_!tpKi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png 1272w, https://substackcdn.com/image/fetch/$s_!tpKi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tpKi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png" width="540" height="72" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:72,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:14382,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!tpKi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png 424w, https://substackcdn.com/image/fetch/$s_!tpKi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png 848w, https://substackcdn.com/image/fetch/$s_!tpKi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png 1272w, https://substackcdn.com/image/fetch/$s_!tpKi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd876f6ac-e4a9-4b85-ba9c-f2e6ffeda78a_540x72.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong>Advantages</strong></p><ul><li><p><strong>Captures long-term dependencies: </strong>LSTM effectively utilizes information from distant time steps, overcoming the vanishing gradient problem in traditional RNNs.</p></li><li><p><strong>Selective memory: </strong>The gating mechanisms allow LSTM to selectively remember or forget information based on relevance, making it robust to noisy inputs.</p></li><li><p><strong>Successful in various applications: </strong>LSTM has achieved state-of-the-art performance in tasks such as language modeling, machine translation, and sentiment analysis.</p></li></ul><p><strong>Limitations</strong></p><ul><li><p><strong>Computational complexity: </strong>LSTM has higher computational complexity due to additional gates and memory cell, resulting in longer training times and increased memory requirements.</p></li><li><p><strong>Parallelization challenges: </strong>The sequential nature of LSTM makes efficient parallelization difficult, limiting the ability to process long sequences simultaneously.</p></li><li><p><strong>Sensitivity to hyperparameters: </strong>LSTM performance can be sensitive to hyperparameter choices, requiring careful tuning and experimentation.</p></li><li><p><strong>Interpretability: </strong>Like other deep learning models, understanding the learned representations and decision-making process in LSTM can be challenging.</p></li></ul><p>Despite these limitations, LSTM remains a powerful and widely used architecture for modeling sequential data and capturing long-term dependencies. Its ability to selectively remember and forget information, along with its robustness to vanishing gradients, has made it a go-to choice for many sequence modeling tasks.</p><p>Variants and extensions of LSTM, such as GRUs and Bidirectional LSTM (BiLSTM), have been proposed to address some of the limitations and improve upon the basic LSTM architecture. These variants aim to simplify the gating mechanisms, reduce computational complexity, or incorporate additional context from both past and future time steps.</p><div><hr></div><h3>GRU</h3><p>GRUs are a type of recurrent neural network architecture that aims to address the vanishing gradient problem and capture long-term dependencies in sequential data. GRUs were introduced as a simpler and more computationally efficient alternative to LSTM networks.</p><p>Like LSTMs, GRUs are designed to handle the challenges of learning long-term dependencies in sequential data. However, GRUs have a simpler structure compared to LSTMs, with fewer parameters and a more streamlined computation process.</p><p>The key idea behind GRUs is to use gating mechanisms to control the flow of information within the network. GRUs have two main gates: the update gate and the reset gate. These gates regulate the amount of information that is retained from the previous time step and the amount of new information that is added at the current time step.</p><p><strong>Architectural Components</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LGIh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LGIh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png 424w, https://substackcdn.com/image/fetch/$s_!LGIh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png 848w, https://substackcdn.com/image/fetch/$s_!LGIh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png 1272w, https://substackcdn.com/image/fetch/$s_!LGIh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LGIh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png" width="809" height="400" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:400,&quot;width&quot;:809,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LGIh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png 424w, https://substackcdn.com/image/fetch/$s_!LGIh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png 848w, https://substackcdn.com/image/fetch/$s_!LGIh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png 1272w, https://substackcdn.com/image/fetch/$s_!LGIh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12da6a8a-4456-4cfe-a424-322cb6af1605_809x400.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure &#8211; GRU architecture</p><p>The GRU architecture consists of the following key components:</p><ol><li><p><strong>Update Gate (z):</strong></p></li></ol><ul><li><p>The update gate determines how much of the previous hidden state should be carried forward to the current time step.</p></li><li><p>It takes the previous hidden state (h<sub>t-1</sub>) and the current input (x<sub>t</sub>) as inputs and produces a value between 0 and 1 for each element in the hidden state.</p></li><li><p>A value of 0 means completely discard the previous hidden state, while a value of 1 means completely retain it.</p></li><li><p>The update gate equation is given by:</p><p></p></li></ul><blockquote><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sWDu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sWDu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png 424w, https://substackcdn.com/image/fetch/$s_!sWDu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png 848w, https://substackcdn.com/image/fetch/$s_!sWDu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png 1272w, https://substackcdn.com/image/fetch/$s_!sWDu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sWDu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png" width="540" height="72" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:72,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:13744,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sWDu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png 424w, https://substackcdn.com/image/fetch/$s_!sWDu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png 848w, https://substackcdn.com/image/fetch/$s_!sWDu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png 1272w, https://substackcdn.com/image/fetch/$s_!sWDu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F066d5f1c-23a8-45da-b6dc-be29b8ff7497_540x72.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Here, W<sub>z</sub> is the weight matrix for the update gate, h<sub>t-1</sub> is the previous hidden state, x<sub>t</sub> is the current input, b<sub>z</sub> is the bias term, and &#963; is the sigmoid activation function.</p></blockquote><ol start="2"><li><p><strong>Reset Gate (r):</strong></p></li></ol><ul><li><p>The reset gate determines how much of the previous hidden state should be forgotten.</p></li><li><p>It takes the previous hidden state (h<sub>t-1</sub>) and the current input (x<sub>t</sub>) as inputs and produces a value between 0 and 1 for each element in the hidden state.</p></li><li><p>A value of 0 means completely forget the previous hidden state, while a value of 1 means completely retain it.</p></li><li><p>The reset gate equation is given by:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pkVO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pkVO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png 424w, https://substackcdn.com/image/fetch/$s_!pkVO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png 848w, https://substackcdn.com/image/fetch/$s_!pkVO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png 1272w, https://substackcdn.com/image/fetch/$s_!pkVO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pkVO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png" width="540" height="66" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c288f6d8-d256-4509-9ced-3daba506508c_540x66.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:66,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:13695,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pkVO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png 424w, https://substackcdn.com/image/fetch/$s_!pkVO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png 848w, https://substackcdn.com/image/fetch/$s_!pkVO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png 1272w, https://substackcdn.com/image/fetch/$s_!pkVO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc288f6d8-d256-4509-9ced-3daba506508c_540x66.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div></li></ul><blockquote><p>Here, W<sub>r</sub> is the weight matrix for the reset gate, and b<sub>r</sub> is the bias term.</p></blockquote><ol start="3"><li><p><strong>Candidate Hidden State (</strong>h<strong>):</strong></p></li></ol><ul><li><p>The candidate hidden state represents the new information that could potentially be added to the current hidden state.</p></li><li><p>It is computed based on the current input (x<sub>t</sub>) and the element-wise product of the reset gate (r<sub>t</sub>) and the previous hidden state (h<sub>t-1</sub>).</p></li><li><p>The candidate hidden state equation is given by:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!o583!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!o583!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png 424w, https://substackcdn.com/image/fetch/$s_!o583!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png 848w, https://substackcdn.com/image/fetch/$s_!o583!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png 1272w, https://substackcdn.com/image/fetch/$s_!o583!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!o583!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png" width="626" height="70" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:70,&quot;width&quot;:626,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:18294,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!o583!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png 424w, https://substackcdn.com/image/fetch/$s_!o583!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png 848w, https://substackcdn.com/image/fetch/$s_!o583!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png 1272w, https://substackcdn.com/image/fetch/$s_!o583!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F394797ee-f6ce-48f0-b5d3-fcbc882a5adf_626x70.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><blockquote><p>Here, W<sub>h</sub> is the weight matrix for the candidate hidden state, b<sub>h</sub> is the bias term, and tanh is the hyperbolic tangent activation function.</p></blockquote><ol start="4"><li><p><strong>Hidden State (h):</strong></p></li></ol><ul><li><p>The hidden state represents the output of the GRU unit at each time step.</p></li><li><p>It is computed as a linear interpolation between the previous hidden state (h<sub>t-1</sub>) and the candidate hidden state (h), controlled by the update gate (z<sub>t</sub>).</p></li><li><p>The hidden state equation is given by:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vh0_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vh0_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png 424w, https://substackcdn.com/image/fetch/$s_!vh0_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png 848w, https://substackcdn.com/image/fetch/$s_!vh0_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png 1272w, https://substackcdn.com/image/fetch/$s_!vh0_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vh0_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png" width="540" height="66" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:66,&quot;width&quot;:540,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:13946,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vh0_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png 424w, https://substackcdn.com/image/fetch/$s_!vh0_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png 848w, https://substackcdn.com/image/fetch/$s_!vh0_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png 1272w, https://substackcdn.com/image/fetch/$s_!vh0_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46b8c6b9-eb85-4b68-85fb-df59f90576ce_540x66.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><blockquote><p>Here, z<sub>t</sub> is the output of the update gate, h<sub>t-1</sub> is the previous hidden state, and ht is the candidate hidden state.</p></blockquote><p><strong>Advantages</strong></p><ul><li><p><strong>Simpler architecture: </strong>GRUs have fewer parameters and a more streamlined computation process compared to LSTMs, leading to faster training and inference times.</p></li><li><p><strong>Efficient capturing of long-term dependencies: </strong>GRUs effectively capture long-term dependencies in sequential data through gating mechanisms.</p></li><li><p><strong>Less prone to overfitting: </strong>The simpler structure and fewer parameters make GRUs less prone to overfitting compared to LSTMs.</p></li><li><p><strong>Good performance in various tasks: </strong>GRUs have shown competitive performance in tasks such as language modeling, machine translation, and speech recognition.</p></li></ul><p><strong>Limitations</strong></p><ul><li><p><strong>Lack of explicit memory cell: </strong>GRUs do not have a separate memory cell like LSTMs, potentially limiting their ability to maintain very long-term dependencies.</p></li><li><p><strong>Sensitivity to hyperparameters: </strong>GRUs can be sensitive to hyperparameter choices, requiring careful tuning for optimal performance.</p></li><li><p><strong>Limited interpretability: </strong>Understanding the learned representations and decision-making process in GRUs can be challenging.</p></li><li><p><strong>Challenges with very long sequences: </strong>GRUs may struggle to capture dependencies in extremely long sequences.</p></li></ul><p>Despite these limitations, GRUs have proven to be a powerful and efficient alternative to LSTMs for modeling sequential data. Their simpler structure and competitive performance have made them a popular choice in various domains, particularly when computational efficiency is a concern.</p><p>Researchers and practitioners often experiment with both LSTMs and GRUs to determine which architecture works best for their specific task and dataset. The choice between LSTMs and GRUs depends on factors such as the complexity of the task, the size of the dataset, the available computational resources, and the desired trade-off between performance and simplicity.</p><div><hr></div><h2>The Transformer Architecture and BERT</h2><p>In recent years, the field of NLP has witnessed significant advancements, largely driven by the development of the Transformer architecture and its variants, such as BERT (Bidirectional Encoder Representations from Transformers). In this section, we will explore the Transformer architecture, its key components, and how BERT builds upon this foundation to achieve state-of-the-art performance on a wide range of NLP tasks.</p><h3><strong>The Transformer Architecture</strong></h3><p>The Transformer architecture, introduced by Vaswani et al. in the paper "Attention Is All You Need," has revolutionized the way NLP models process and understand sequential data. Unlike previous architectures, such as RNNs and LSTM networks, the Transformer relies solely on attention mechanisms to capture dependencies between input tokens.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EtTh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EtTh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png 424w, https://substackcdn.com/image/fetch/$s_!EtTh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png 848w, https://substackcdn.com/image/fetch/$s_!EtTh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png 1272w, https://substackcdn.com/image/fetch/$s_!EtTh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EtTh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png" width="805" height="760" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:760,&quot;width&quot;:805,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EtTh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png 424w, https://substackcdn.com/image/fetch/$s_!EtTh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png 848w, https://substackcdn.com/image/fetch/$s_!EtTh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png 1272w, https://substackcdn.com/image/fetch/$s_!EtTh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e36c3e2-08e3-43b0-989e-c30752bb7b20_805x760.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure &#8211; Transformer architecture</p><p><strong>Encoder-Decoder Structure</strong></p><p>The Transformer consists of an encoder and a decoder, each composed of a stack of identical layers. The encoder takes the input sequence and generates a contextualized representation, while the decoder generates the output sequence based on the encoder's output and the previously generated tokens.</p><p><strong>Multi-Head Attention</strong></p><p>At the core of the Transformer architecture is the multi-head attention mechanism. Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions. It enables the model to capture complex relationships and dependencies between tokens in the input sequence.</p><p>In multi-head attention, the input sequence is projected into multiple query, key, and value vectors. The attention weights are computed by taking the dot product between the query and key vectors, followed by a softmax function. The output is obtained by multiplying the attention weights with the value vectors.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8SoO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8SoO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png 424w, https://substackcdn.com/image/fetch/$s_!8SoO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png 848w, https://substackcdn.com/image/fetch/$s_!8SoO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png 1272w, https://substackcdn.com/image/fetch/$s_!8SoO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8SoO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png" width="986" height="566" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:566,&quot;width&quot;:986,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8SoO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png 424w, https://substackcdn.com/image/fetch/$s_!8SoO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png 848w, https://substackcdn.com/image/fetch/$s_!8SoO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png 1272w, https://substackcdn.com/image/fetch/$s_!8SoO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F27bb1a7d-c202-4462-b7d8-40fa66e44133_986x566.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure  &#8211; Multi-Head attention</p><p><strong>Position-wise Feed-Forward Networks</strong></p><p>In addition to multi-head attention, each layer in the Transformer also includes position-wise feed-forward networks. These networks are applied independently to each position in the sequence and consist of two linear transformations with a ReLU activation in between. The feed-forward networks help the model learn higher-level representations and capture complex patterns in the input.</p><p><strong>Positional Encoding</strong></p><p>Since the Transformer does not rely on recurrent connections, it needs a way to incorporate positional information into the input representations. This is achieved through positional encoding, where each position in the sequence is assigned a unique vector that encodes its relative position. The positional encodings are added to the input embeddings, allowing the model to capture the order and relative positions of the tokens.</p><div><hr></div><h3><strong>BERT: Bidirectional Encoder Representations from Transformers</strong></h3><p>Building upon the Transformer architecture, BERT has become one of the most influential and widely-used models in NLP. BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XAfJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XAfJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png 424w, https://substackcdn.com/image/fetch/$s_!XAfJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png 848w, https://substackcdn.com/image/fetch/$s_!XAfJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png 1272w, https://substackcdn.com/image/fetch/$s_!XAfJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XAfJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png" width="891" height="751" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:751,&quot;width&quot;:891,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XAfJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png 424w, https://substackcdn.com/image/fetch/$s_!XAfJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png 848w, https://substackcdn.com/image/fetch/$s_!XAfJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png 1272w, https://substackcdn.com/image/fetch/$s_!XAfJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3c83dd-9092-49c4-b52b-f275327999c4_891x751.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure &#8211; BERT architecture</p><p><strong>BERT is pre-trained on two unsupervised tasks</strong></p><ul><li><p><strong>Masked Language Modeling (MLM):</strong> In this task, a random subset of tokens in the input sequence is masked, and the objective is to predict the original vocabulary ID of the masked word based only on its context. This allows the model to learn bidirectional representations by considering both the left and right context.</p></li><li><p><strong>Next Sentence Prediction (NSP):</strong> In this task, the model is given a pair of sentences and learns to predict whether the second sentence follows the first sentence in the original text. This helps the model understand the relationship between sentences, which is crucial for tasks like question answering and natural language inference.</p></li></ul><p><strong>Input Representation</strong></p><p>BERT takes a sequence of tokens as input, which can be a single sentence or a pair of sentences separated by a special token ([SEP]). The input representation for each token is constructed by summing the corresponding token embedding, segment embedding, and position embedding.</p><p><strong>Fine-tuning</strong></p><p>One of the key strengths of BERT is its ability to be fine-tuned for specific downstream tasks with minimal modifications. By adding a task-specific output layer on top of the pre-trained BERT model, it can be adapted to various NLP tasks, such as sentiment analysis, named entity recognition, and question answering.</p><p>During fine-tuning, the pre-trained BERT parameters are initialized, and the model is trained on a labeled dataset specific to the downstream task. The model learns to map the input sequences to the desired output format, leveraging the rich representations learned during pre-training.</p><p><strong>Advantages and Limitations</strong></p><p>BERT has several advantages that have contributed to its success:</p><ul><li><p><strong>Bidirectional representations:</strong> BERT's ability to learn from both left and right context allows it to capture more accurate and nuanced representations of words and their relationships.</p></li><li><p><strong>Unsupervised pre-training:</strong> By pre-training on large unlabeled text corpora, BERT can learn general language representations that can be fine-tuned for specific tasks with relatively small labeled datasets.</p></li><li><p><strong>State-of-the-art performance:</strong> BERT has achieved state-of-the-art results on a wide range of NLP benchmarks, demonstrating its effectiveness in capturing contextual information and learning rich representations.</p></li></ul><p>However, <strong>BERT also has some limitations:</strong></p><ul><li><p><strong>Computational complexity:</strong> BERT has a large number of parameters and requires significant computational resources for pre-training and fine-tuning, which can be a challenge for resource-constrained environments.</p></li><li><p><strong>Limited sequence length:</strong> BERT has a fixed maximum sequence length (typically 512 tokens), which can be a limitation for tasks that require processing longer sequences or documents.</p></li><li><p><strong>Lack of interpretability:</strong> Like many deep learning models, BERT's internal representations and decision-making process can be difficult to interpret, which can be a concern in certain applications where explainability is important.</p></li></ul><p>Despite these limitations, BERT has had a profound impact on the field of NLP and has paved the way for further advancements in language modeling and understanding.</p><p><strong>Code:</strong></p><pre><code>!pip install transformers

import numpy as np

import pandas as pd

from sklearn.model_selection import train_test_split

from sklearn.svm import SVC

from sklearn.model_selection import cross_val_score

import torch

import transformers as ppb

df = pd.read_csv('/path/train.tsv', delimiter='\t', header=None)

batch_1 = df[:2500]

# For DistilBERT:

model_class, tokenizer_class, pretrained_weights = (ppb.DistilBertModel, ppb.DistilBertTokenizer, 'distilbert-base-uncased')

## Want BERT instead of distilBERT? Uncomment the following line:

#model_class, tokenizer_class, pretrained_weights = (ppb.BertModel, ppb.BertTokenizer, 'bert-base-uncased')

# Load pretrained model/tokenizer

tokenizer = tokenizer_class.from_pretrained(pretrained_weights)

model = model_class.from_pretrained(pretrained_weights)

tokenized = batch_1[0].apply((lambda x: tokenizer.encode(x, add_special_tokens=True)))

max_len = 0

for i in tokenized.values:

if len(i) &gt; max_len:

max_len = len(i)

padded = np.array([i + [0]*(max_len-len(i)) for i in tokenized.values])

attention_mask = np.where(padded != 0, 1, 0)

attention_mask.shape

input_ids = torch.tensor(padded)

attention_mask = torch.tensor(attention_mask)

with torch.no_grad():

last_hidden_states = model(input_ids, attention_mask=attention_mask)

features = last_hidden_states[0][:,0,:].numpy()

labels = batch_1[1]

train_features, test_features, train_labels, test_labels = train_test_split(features, labels)

lr_clf = SVC()

lr_clf.fit(train_features, train_labels)

lr_clf.score(test_features, test_labels)</code></pre><p><strong>Description:</strong></p><ul><li><p>The code demonstrates how to use the BERT model for sentiment classification of movie reviews using the Hugging Face Transformers library. The dataset consists of movie reviews labeled as either positive or negative. Here's a step-by-step explanation:</p></li><li><p>The necessary libraries are imported, including NumPy, pandas, scikit-learn, PyTorch, and the Transformers library.</p></li><li><p>The movie review dataset is loaded from a TSV file using pandas. In this example, only the first 2500 rows of the dataset are used (stored in batch_1).</p></li><li><p>The BERT model and tokenizer are loaded using the BertModel and BertTokenizer classes from the Transformers library. The 'bert-base-uncased' pre-trained weights are used.</p></li><li><p>An attention mask is created to indicate which tokens are actual words (1) and which are padding tokens (0).</p></li><li><p>The padded input sequences and attention mask are converted to PyTorch tensors.</p></li><li><p>The BERT model is used to generate hidden state representations for the input sequences. The last hidden states are obtained by passing the input IDs and attention mask to the model.</p></li><li><p>The features are extracted from the last hidden states by taking the first token's representation (usually the [CLS] token) for each sequence.</p></li><li><p>The sentiment labels (positive or negative) are obtained from batch_1[1].</p></li><li><p>A Support Vector Machine (SVM) classifier is initialized and trained on the training features and labels.</p><div><hr></div></li></ul><h1><strong>Connect with Me</strong></h1><ol><li><p>If you have any inquiries, feel free to reach out via message or email.</p></li></ol><blockquote><p><em><strong><a href="https://abonia1.github.io/">Website/Newletter</a></strong></em></p><p><em>Connect with me on<strong> <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p><p><em>Find me on<strong> <a href="https://github.com/Abonia1">Github</a></strong></em></p><p><em>Visit my technical channel on <strong><a href="https://www.youtube.com/@AboniaSojasingarayar">Youtube</a></strong></em></p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Chapter 2 - Traditional and Modern Text Representation Techniques]]></title><description><![CDATA[Exploring foundational and advanced methods for representing textual data]]></description><link>https://aboniasojasingarayar.substack.com/p/chapter-2-traditional-and-modern</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/chapter-2-traditional-and-modern</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Mon, 27 Jan 2025 09:01:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!j79V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>Table of Content</strong></p><ul><li><p>2.1 Text Representation Techniques: Bag-of-words, TF-IDF, Word2Vec, GloVe, FastText</p><ul><li><p>2.1.1 Mathematical foundations of text representation techniques</p></li><li><p>2.1.2 Algorithmic implementation details</p></li><li><p>2.1.3 Comparative evaluation of different techniques</p></li></ul></li><li><p>2.2 Embeddings (ELMo, BERT) and contextual representation</p><ul><li><p>2.2.1 Architecture and design of ELMo and BERT</p></li><li><p>2.2.2 Pre-Training approaches for contextual embeddings</p></li><li><p>2.2.3 Fine-tuning contextual embeddings</p><p></p><div><hr></div></li></ul></li></ul><p>Written in Collaboration - Special thanks to <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Manan Thakkar&quot;,&quot;id&quot;:300408543,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10e24dac-fd32-4234-ac63-24a90d0f362c_144x144.png&quot;,&quot;uuid&quot;:&quot;9625a19d-ab98-43f0-a382-8a5b969b2474&quot;}" data-component-name="MentionToDOM"></span> for his valuable contributions to the structure and content of this chapter, shaping it into an informative and accessible resource for understanding text representation techniques.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The ability to understand and generate human language is central to NLP. To achieve this, NLP systems must be able to represent the meaning of language in a way that computers can process. This chapter provides an overview of key techniques for representing text as mathematical vectors, enabling statistical and neural network models to extract semantic information from language data.</p><p>We first introduce the bag-of-words model and term frequency-inverse document frequency (TF-IDF), classic approaches that represent text based on word frequencies. Then, we explore neural embedding techniques like Word2Vec, GloVe, and FastText that capture semantic relationships between words. Next, we discuss contextual representation techniques like ELMo and BERT that incorporate the context around words to represent meaning.</p><p>By the end of this chapter, you will have strong conceptual and mathematical foundations regarding text representation, enabling you to develop predictive NLP systems. The techniques presented form the basis for contemporary NLP with deep neural networks.</p><p>In this chapter, we will cover the following topics:</p><ul><li><p>Text Representation Techniques: Mathematical foundations, algorithms, and evaluation for bag-of-words, TF-IDF, Word2Vec, GloVe, and FastText</p></li><li><p>Contextualized Embeddings: Architectures, pre-training, and fine-tuning methods for models like ELMo and BERT</p><div><hr></div></li></ul><h2>Text Representation Techniques: Bag-of-words, TF-IDF, Word2Vec, GloVe, <strong>FastText</strong></h2><p>Text representation is a fundamental task in NLP that transforms raw text into vector representations with numerical values that can be readily used by machine learning models. Early techniques like bag-of-words and TF-IDF represented text based on word frequencies, ignoring word order but capturing basic statistical patterns. More recent neural embedding methods like Word2Vec, GloVe, and FastText aim to encode semantic similarity between words into vector representations. In this section, we provide an overview of these major approaches for representing text, outlining the key concepts, algorithms, and evaluation procedures. We begin with a discussion of bag-of-words and TF-IDF, building up intuitions before presenting more complex semantic embedding strategies. By the end, you will have strong foundations regarding how NLP systems numerically represent language for prediction and analysis using vector representations.</p><div><hr></div><h3><strong>Bag-of-words</strong></h3><p>The bag-of-words model is one of the earliest and simplest ways to represent text numerically for machine learning. As the name suggests, it treats each document as an unordered collection of words, disregarding grammar and word order entirely. The central idea behind this representation is to capture the basic statistical information about the frequencies of words within a document. Each document is modeled as a multiset or "bag" of its words. The grammar and ordering of the words is ignored, and only the word counts matter.</p><p>Concretely, each document is represented as a vector containing the counts for each unique word or token in the vocabulary. This vector encodes the word frequencies but disregards contextual relationships between words. The dimensionality of the vector is equal to the number of unique words across the corpus.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!j79V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!j79V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png 424w, https://substackcdn.com/image/fetch/$s_!j79V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png 848w, https://substackcdn.com/image/fetch/$s_!j79V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png 1272w, https://substackcdn.com/image/fetch/$s_!j79V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!j79V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png" width="1168" height="432" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d6fdc684-068a-4c44-9564-952dc0816608_1168x432.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:432,&quot;width&quot;:1168,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115677,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!j79V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png 424w, https://substackcdn.com/image/fetch/$s_!j79V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png 848w, https://substackcdn.com/image/fetch/$s_!j79V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png 1272w, https://substackcdn.com/image/fetch/$s_!j79V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6fdc684-068a-4c44-9564-952dc0816608_1168x432.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Example:</strong></p><p>To illustrate how the bag-of-words model transforms text into numerical vectors, let's walk through a simple example. Consider the following two sentences:</p><p><strong>Document 1:</strong> &#8220;NLP has advanced rapidly in recent years&#8221;</p><p><strong>Document 2:</strong> &#8220;LLM models have transformed NLP systems and NLP capabilities&#8221;</p><p>We first preprocess the sentences by lowercasing, removing stopwords, and lemmatizing to obtain:</p><p>Corpus = ['nlp advance rapidly recent years',</p><p>'llm model transform nlp systems nlp capabilities']</p><p>Looking at all the unique words, our vocabulary is:</p><p>['advance', 'capabilities', 'llm', 'model', 'nlp', 'rapidly', 'recent', 'systems', 'transform', 'years']</p><p>With this 10-word vocabulary, we can represent each sentence as a 10-dimensional count vector. Element i contains the count of the i-th vocabulary word in that sentence.</p><p>For sentence 1, the count of words is as follow:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oOns!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oOns!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png 424w, https://substackcdn.com/image/fetch/$s_!oOns!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png 848w, https://substackcdn.com/image/fetch/$s_!oOns!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png 1272w, https://substackcdn.com/image/fetch/$s_!oOns!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oOns!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png" width="570" height="438" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:438,&quot;width&quot;:570,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:28816,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oOns!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png 424w, https://substackcdn.com/image/fetch/$s_!oOns!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png 848w, https://substackcdn.com/image/fetch/$s_!oOns!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png 1272w, https://substackcdn.com/image/fetch/$s_!oOns!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a53ec2f-9a31-4224-9edd-1c44c19942b7_570x438.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>For Sentence 1, the vector is [1, 0, 0, 0, 1, 1, 1, 0, 0, 1], with counts for "nlp", "advance", "rapidly", "recent", and "years".</p><p>Now for sentence 2, the scoring would be like</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wMx9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wMx9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png 424w, https://substackcdn.com/image/fetch/$s_!wMx9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png 848w, https://substackcdn.com/image/fetch/$s_!wMx9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png 1272w, https://substackcdn.com/image/fetch/$s_!wMx9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wMx9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png" width="570" height="438" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:438,&quot;width&quot;:570,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:28988,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wMx9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png 424w, https://substackcdn.com/image/fetch/$s_!wMx9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png 848w, https://substackcdn.com/image/fetch/$s_!wMx9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png 1272w, https://substackcdn.com/image/fetch/$s_!wMx9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a7ed01c-9b74-47fd-a217-16b9456769c6_570x438.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Similarly, the vector for Sentence 2 is [0, 1, 1, 1, 2, 0, 0, 1, 1, 0], with counts for the rest of the vocabulary words.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KxgV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KxgV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png 424w, https://substackcdn.com/image/fetch/$s_!KxgV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png 848w, https://substackcdn.com/image/fetch/$s_!KxgV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png 1272w, https://substackcdn.com/image/fetch/$s_!KxgV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KxgV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png" width="1394" height="252" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:252,&quot;width&quot;:1394,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:36738,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KxgV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png 424w, https://substackcdn.com/image/fetch/$s_!KxgV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png 848w, https://substackcdn.com/image/fetch/$s_!KxgV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png 1272w, https://substackcdn.com/image/fetch/$s_!KxgV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff48e91e0-9f5d-483e-a562-713ae11e2e08_1394x252.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>The approach used in example two is the one that is generally used in the Bag-of-Words technique, the reason being that the datasets used in Machine learning are tremendously large and can contain vocabulary of a few thousand or even millions of words. Hence, preprocessing the text before using bag-of-words is a better way to go.</p><h5><strong>Advantages</strong></h5><ul><li><p><strong>Simplicity:</strong> The bag-of-words approach is very simple to understand and implement. Tokenization and counting frequencies are computationally straightforward.</p></li><li><p><strong>Efficiency:</strong> Constructing the vocabulary and transforming documents into count vectors can be done very efficiently with sparse matrix representations. This allows bag-of-words to scale to large datasets.</p></li><li><p><strong>Effectiveness for classification:</strong> Despite its flaws, bag-of-words works reasonably well for document classification tasks. The word frequencies provide enough statistical signal for simple categorical predictions.</p></li></ul><h5><strong>Limitations</strong></h5><ul><li><p><strong>Lacks word order:</strong> Completely discarding word order is a major limitation. The model cannot differentiate "dogs bite men" from "men bite dogs" as the counts are identical. Position and sequence are lost.</p></li><li><p><strong>Ignores context:</strong> Counts are aggregated globally across the document, ignoring local context. The distributions of words in different sections of text are not captured.</p></li><li><p><strong>No phrase modeling:</strong> Bag-of-words looks at single words only, not phrases. So "New York Times" is treated as separate unrelated words.</p></li><li><p><strong>High dimensionality:</strong> The vocabulary size directly determines the number of dimensions. Real-world datasets can easily have vocabularies in the millions, leading to extremely high-dimensional representations. This causes sparsity and statistical issues.</p></li><li><p><strong>No semantics:</strong> Synonyms like "smart" and "intelligent" are treated as unrelated. Semantic meaning is lost without representations of word-word relations.</p></li><li><p><strong>Rare words:</strong> Words not seen during training are ignored entirely. But rare or new words can be highly informative, such as technical terms or neologisms.</p></li></ul><p><strong>Here is sample Python code to generate a bag-of-words representation:</strong></p><pre><code>import numpy as np

corpus = [' nlp advance rapidly recent years ',

' llm model transform nlp systems nlp capabilities ']

vocab = ['advance', 'capabilities', 'llm', 'model', 'nlp', 'rapidly', 'recent', 'systems', 'transform', 'years']

position = {}

for i, token in enumerate(vocab):

position[token] = i

print(position)

bow_matrix = np.zeros((len(preprocessed_corpus), len(vocab)))

for i, preprocessed_sentence in enumerate(preprocessed_corpus):

for token in preprocessed_sentence.split():

bow_matrix[i][position[token]] = bow_matrix[i][position[token]] + 1

bow_matrix</code></pre><p><strong>OUTPUT:</strong></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HBB-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HBB-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png 424w, https://substackcdn.com/image/fetch/$s_!HBB-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png 848w, https://substackcdn.com/image/fetch/$s_!HBB-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png 1272w, https://substackcdn.com/image/fetch/$s_!HBB-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HBB-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png" width="534" height="134" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:134,&quot;width&quot;:534,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:21871,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HBB-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png 424w, https://substackcdn.com/image/fetch/$s_!HBB-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png 848w, https://substackcdn.com/image/fetch/$s_!HBB-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png 1272w, https://substackcdn.com/image/fetch/$s_!HBB-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c2e7255-076c-45ad-91dd-e497eea0deaf_534x134.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>This code snippet demonstrates how to construct a bag-of-words representation matrix from a preprocessed text corpus.</p><p>First, the pre-processed corpus and vocabulary as a list of all unique words extracted from the corpus using word tokenizer are defined. Then, a position dictionary is created to map each word to an index value.</p><p>Next, a blank bag-of-words matrix is initialized with shape (num_documents, vocabulary_size), to store the word counts for each document. We loop through each preprocessed sentence in the corpus, splitting into words. For each word, we increment the count at the corresponding vocabulary index in the bow_matrix for that document.</p><p>By the end, bow_matrix will contain the bag-of-words representation where each row is the word count vector for a document. This matrix encodes the word frequencies but lacks word order and semantics.</p><p>This code demonstrates an efficient and vectorized approach to generate bag-of-words representations from text corpora before feeding into machine learning models. The simple counts capture basic statistics but not linguistic meaning.</p><div><hr></div><h3><strong>TF-IDF</strong></h3><p>The standard bag-of-words representation that underpins many text analysis models utilizes raw term frequencies (TF) as the primary feature values. However, frequency alone does not always capture relevance. Extremely common words can appear with high regularity just by chance, overwhelming the signal from rarer but more informative keywords.</p><p>To counteract this tendency, the TF-IDF scheme introduces a second weighting factor called inverse document frequency (IDF). This rebalances weights so that ubiquitous terms are scaled down and distinctive salient terms are scaled up. Specifically, IDF measures how common or unique a word is across the entire document corpus by taking the logarithm of the inverse fraction of documents containing that word. Values are higher for words occurring in fewer documents, tapering for broadly common words.</p><p>Multiplying the TF and IDF assigns the highest scores to words that strike an optimal balance - frequent enough locally to be relevant for a document, but globally rare enough to uniquely describe a document's meaning. This amplifies keywords that characterize document content rather than broadly useless words.</p><p>For example, &#8220;phone&#8221; may appear often in a cell phone manual, but so do generic terms like "settings" and "menu" across many documents. The IDF downweights these widely common words so that meaningful keywords like &#8220;phone&#8221; rise to the top.</p><p>The TF-IDF score for a term in a document is calculated as the product of its TF and IDF. Here are the equations for TF and IDF, and the overall TF-IDF:</p><ul><li><p>Term-Frequency (TF):</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FnAk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FnAk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png 424w, https://substackcdn.com/image/fetch/$s_!FnAk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png 848w, https://substackcdn.com/image/fetch/$s_!FnAk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png 1272w, https://substackcdn.com/image/fetch/$s_!FnAk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FnAk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png" width="1156" height="122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a367594-d0af-4287-b689-485b329f656b_1156x122.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:122,&quot;width&quot;:1156,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:49570,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FnAk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png 424w, https://substackcdn.com/image/fetch/$s_!FnAk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png 848w, https://substackcdn.com/image/fetch/$s_!FnAk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png 1272w, https://substackcdn.com/image/fetch/$s_!FnAk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a367594-d0af-4287-b689-485b329f656b_1156x122.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><ul><li><p>Inverse Document Frequency (IDF):</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2lO-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2lO-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png 424w, https://substackcdn.com/image/fetch/$s_!2lO-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png 848w, https://substackcdn.com/image/fetch/$s_!2lO-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png 1272w, https://substackcdn.com/image/fetch/$s_!2lO-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2lO-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png" width="1256" height="134" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6d831533-e706-4843-bf21-60be85363d1f_1256x134.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:134,&quot;width&quot;:1256,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:54474,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2lO-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png 424w, https://substackcdn.com/image/fetch/$s_!2lO-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png 848w, https://substackcdn.com/image/fetch/$s_!2lO-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png 1272w, https://substackcdn.com/image/fetch/$s_!2lO-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d831533-e706-4843-bf21-60be85363d1f_1256x134.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><ul><li><p>TF-IDF:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lHi9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lHi9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png 424w, https://substackcdn.com/image/fetch/$s_!lHi9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png 848w, https://substackcdn.com/image/fetch/$s_!lHi9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png 1272w, https://substackcdn.com/image/fetch/$s_!lHi9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lHi9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png" width="1256" height="134" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b2382264-7635-440f-856f-31578ceea69c_1256x134.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:134,&quot;width&quot;:1256,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:25115,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lHi9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png 424w, https://substackcdn.com/image/fetch/$s_!lHi9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png 848w, https://substackcdn.com/image/fetch/$s_!lHi9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png 1272w, https://substackcdn.com/image/fetch/$s_!lHi9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2382264-7635-440f-856f-31578ceea69c_1256x134.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong>Example:</strong></p><p>To illustrate how the bag-of-words model transforms text into numerical vectors, let's walk through a simple example. Consider the following two sentences:</p><p><strong>Document 1:</strong> &#8220;NLP has advanced rapidly in recent years&#8221;</p><p><strong>Document 2:</strong> &#8220;LLM models have transformed NLP systems and NLP capabilities&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HGXE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HGXE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png 424w, https://substackcdn.com/image/fetch/$s_!HGXE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png 848w, https://substackcdn.com/image/fetch/$s_!HGXE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png 1272w, https://substackcdn.com/image/fetch/$s_!HGXE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HGXE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png" width="1250" height="1126" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1126,&quot;width&quot;:1250,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:349280,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HGXE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png 424w, https://substackcdn.com/image/fetch/$s_!HGXE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png 848w, https://substackcdn.com/image/fetch/$s_!HGXE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png 1272w, https://substackcdn.com/image/fetch/$s_!HGXE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d90c05c-5baf-4130-849b-dbae9efa9417_1250x1126.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p><strong>Example:</strong></p><p>To illustrate TF-IDF vectorization, consider the same pre-processed statements as our corpus:</p><p><strong>Document 1:</strong> &#8220;nlp advance rapidly recent years&#8221;</p><p><strong>Document 2:</strong> &#8220;llm model transform nlp systems nlp capabilities&#8221;</p><p>The full set of unique words across both documents forms the vocabulary V:</p><p>V = {nlp, advance, rapidly, recent, years, llm, model, transform, systems, capabilities}</p><p>Total Vocabulary Size = 10 terms</p><p><strong>Calculate Term Frequencies (TF)</strong></p><p><strong>Doc 1:</strong></p><ul><li><p>Number of words/terms = 5 terms</p></li><li><p>Calculate TF per term as: Number of occurrences / Total terms</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Rs9p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Rs9p!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png 424w, https://substackcdn.com/image/fetch/$s_!Rs9p!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png 848w, https://substackcdn.com/image/fetch/$s_!Rs9p!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png 1272w, https://substackcdn.com/image/fetch/$s_!Rs9p!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Rs9p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png" width="586" height="448" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:448,&quot;width&quot;:586,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:46987,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Rs9p!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png 424w, https://substackcdn.com/image/fetch/$s_!Rs9p!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png 848w, https://substackcdn.com/image/fetch/$s_!Rs9p!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png 1272w, https://substackcdn.com/image/fetch/$s_!Rs9p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa617f839-c255-4d07-b1c0-e6e79d4fa039_586x448.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Doc 2:</strong></p><ul><li><p>Number of words/terms = 7 terms</p></li><li><p>Repeat TF calculation:</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6G29!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6G29!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png 424w, https://substackcdn.com/image/fetch/$s_!6G29!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png 848w, https://substackcdn.com/image/fetch/$s_!6G29!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png 1272w, https://substackcdn.com/image/fetch/$s_!6G29!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6G29!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png" width="582" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98aa6ccd-294d-4382-86f2-c51862386273_582x745.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:582,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:66877,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6G29!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png 424w, https://substackcdn.com/image/fetch/$s_!6G29!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png 848w, https://substackcdn.com/image/fetch/$s_!6G29!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png 1272w, https://substackcdn.com/image/fetch/$s_!6G29!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98aa6ccd-294d-4382-86f2-c51862386273_582x745.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Calculate the Document Frequency (DF) and Inverse Document Frequencies (IDF)</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wf9-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wf9-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png 424w, https://substackcdn.com/image/fetch/$s_!wf9-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png 848w, https://substackcdn.com/image/fetch/$s_!wf9-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png 1272w, https://substackcdn.com/image/fetch/$s_!wf9-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wf9-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png" width="730" height="850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/255d6cba-1d7d-487e-99d7-775564801f77_730x850.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:850,&quot;width&quot;:730,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:125834,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wf9-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png 424w, https://substackcdn.com/image/fetch/$s_!wf9-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png 848w, https://substackcdn.com/image/fetch/$s_!wf9-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png 1272w, https://substackcdn.com/image/fetch/$s_!wf9-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F255d6cba-1d7d-487e-99d7-775564801f77_730x850.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Calculate TF-IDF Scores</strong></p><p>TF-IDF = TF * IDF</p><p>Showing vector scores:</p><p><code>Doc 1: [0, 0.06020600, 0.06020600, 0.06020600, 0.06020600, 0, 0, 0, 0, 0]</code></p><p><code>Doc 2: [0, 0, 0, 0, 0, 0.0427665, 0.0427665, 0.0427665, 0.0427665, 0.0427665]</code></p><h5><strong>Advantages</strong></h5><ul><li><p><strong>Measures Term Relevance:</strong> TF-IDF effectively identifies terms that are relevant in a particular document compared to the entire corpus based on frequency contrasts. High scores mean a term captures key topical content.</p></li><li><p><strong>Handles Large Corpora:</strong> The TF-IDF formulation scales to huge text collections with sparse matrix representations, making it suitable for big data scenarios. Memory-efficient implementations are feasible.</p></li><li><p><strong>Reduces Impact of Common Terms:</strong> TF-IDF automatically downweights ubiquitous stop words that lack informational value, preventing domination by generic frequent terms.</p></li><li><p><strong>Applicable for Many Tasks:</strong> TF-IDF creates descriptive feature representations that provide good performance across applications like search, classification, clustering, and more.</p></li><li><p><strong>Interpretable Scores:</strong> The TF-IDF scores assign direct quantitative assessments of term importance. This enables qualitative analysis of a document's key themes.</p></li></ul><h5><strong>Limitations</strong></h5><ul><li><p><strong>No Semantic Analysis:</strong> TF-IDF treats semantic relationships between words superficially. Synonyms and related terms are viewed as independent.</p></li><li><p><strong>Assumption of Term Independence:</strong> The formulation assumes words occurring in a document are independent, whereas natural language contains more complex linguistic relationships.</p></li><li><p><strong>High Dimensionality:</strong> The vocabulary size can easily reach millions, leading to very wide, sparse matrices that create statistical issues.</p></li><li><p><strong>Lacks Word Order Modeling:</strong> All terms are represented equally regardless of sequence, discarding order dependence and syntactic context.</p></li><li><p><strong>Hyper parameters Performance Impact:</strong> Stopword removal and smoothing schemes for IDF can significantly sway overall results. TF-IDF requires careful tuning.</p></li></ul><p><strong>Here is sample Python code to generate a TF-IDF vector representation:</strong></p><pre><code>import numpy as np

corpus = [' nlp advance rapidly recent years ',

' llm model transform nlp systems nlp capabilities ']

vectorizer = TfidfVectorizer()

tf_idf_matrix = vectorizer.fit_transform(preprocessed_corpus)

print(vectorizer.get_feature_names_out())

print(tf_idf_matrix.toarray())

print("\nThe shape of the TF-IDF matrix is: ", tf_idf_matrix.shape)</code></pre><p><strong>OUTPUT:</strong></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GZJW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GZJW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png 424w, https://substackcdn.com/image/fetch/$s_!GZJW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png 848w, https://substackcdn.com/image/fetch/$s_!GZJW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png 1272w, https://substackcdn.com/image/fetch/$s_!GZJW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GZJW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png" width="1024" height="254" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:254,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GZJW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png 424w, https://substackcdn.com/image/fetch/$s_!GZJW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png 848w, https://substackcdn.com/image/fetch/$s_!GZJW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png 1272w, https://substackcdn.com/image/fetch/$s_!GZJW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26e772-3ea5-48c9-a8ba-7c9fa6f5e6d8_1024x254.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><div><hr></div><h3><strong>Word2Vec</strong></h3><p>Word2Vec is a popular technique for learning word embeddings, which are dense vector representations of words in a continuous vector space. It was introduced by Tomas Mikolov and his colleagues at Google in 2013. The main idea behind Word2Vec is to capture semantic and syntactic relationships between words based on their co-occurrence patterns in a large corpus of text.</p><p><strong>The Need for Word2Vec:</strong></p><p>Traditional bag-of-words and TF-IDF representations have several limitations. They treat words as discrete and independent entities, ignoring their semantic relationships and context. This makes it difficult to capture the meaning and similarity between words. Word2Vec addresses these limitations by learning dense vector representations that capture the semantic and syntactic relationships between words.</p><p><strong>Techniques in Word2Vec:</strong></p><p>Word2Vec consists of two main techniques for learning word embeddings:</p><p><strong>1.Continuous Bag-of-Words (CBOW):</strong></p><ul><li><p>In CBOW, the model predicts the target word based on its surrounding context words.</p></li><li><p>The input to the model is the context words (e.g., a window of words around the target word), and the output is the target word.</p></li><li><p>The objective is to maximize the probability of predicting the target word given its context.</p></li><li><p>CBOW is computationally efficient and works well with small datasets.</p></li></ul><p><strong>2.Skip-Gram:</strong></p><ul><li><p>In Skip-Gram, the model predicts the surrounding context words given a target word.</p></li><li><p>The input to the model is the target word, and the output is the context words.</p></li><li><p>The objective is to maximize the probability of predicting the context words given the target word.</p></li><li><p>Skip-Gram is more effective in capturing rare words and performs better with larger datasets.</p></li></ul><p><strong>Importance of Word2Vec:</strong></p><p>Word2Vec has revolutionized the field of NLP by providing a powerful way to represent words as dense vectors. Its importance can be highlighted through the following points:</p><ul><li><p><strong>Semantic Relationships:</strong> Word2Vec captures semantic relationships between words. Words with similar meanings are mapped to nearby points in the vector space. This allows for measuring the semantic similarity between words using cosine similarity or Euclidean distance.</p></li><li><p><strong>Analogy Reasoning:</strong> Word2Vec enables analogy reasoning through simple vector arithmetic. For example, the analogy "king - man + woman = queen" can be solved by performing vector operations on the corresponding word embeddings.</p></li><li><p><strong>Transfer Learning:</strong> Word embeddings learned using Word2Vec can be used as pre-trained features for various downstream NLP tasks, such as sentiment analysis, named entity recognition, and text classification. This transfer learning approach has shown significant improvements in performance.</p></li><li><p><strong>Dimensionality Reduction:</strong> Word2Vec reduces the dimensionality of word representations compared to one-hot encoding or TF-IDF. This makes it more computationally efficient and allows for better generalization.</p></li><li><p><strong>Language Model Pre-training:</strong> Word2Vec has paved the way for more advanced language model pre-training techniques, such as GloVe, FastText, and contextual embedding like ELMo and BERT. These models build upon the ideas of Word2Vec to capture even richer linguistic information.</p></li></ul><p><strong>Skip-Gram:</strong></p><p>The Skip-Gram model is a powerful technique for learning dense vector representations of words, known as word embeddings. It aims to predict the surrounding context words given a target word, based on the idea that words occurring in similar contexts tend to have similar meanings.</p><p>To understand how the Skip-Gram model works, let's consider an example sentence:</p><p>"NLP has advanced rapidly with the rise of LLM models."</p><p>In the Skip-Gram model, we choose a target word and aim to predict its surrounding context words within a specified window size. Let's consider the target word "LLM" and a window size of 2.</p><p><strong>Target word:</strong> "LLM"</p><p><strong>Context words:</strong> "the", "rise", "of", "models"</p><p>The Skip-Gram model will generate pairs of (target word, context word) as training examples:</p><ul><li><p>(LLM, the)</p></li><li><p>(LLM, rise)</p></li><li><p>(LLM, of)</p></li><li><p>(LLM, models)</p></li></ul><p>These pairs are used as input to a neural network. The input to the network is a one-hot encoded vector representing the target word, and the output is a one-hot encoded vector representing the context word. The objective of the model is to maximize the probability of predicting the context words given the target word. Let's explore these components step by step, using the example sentence: "NLP has advanced rapidly with the rise of LLM models."</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CX1k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CX1k!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png 424w, https://substackcdn.com/image/fetch/$s_!CX1k!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png 848w, https://substackcdn.com/image/fetch/$s_!CX1k!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png 1272w, https://substackcdn.com/image/fetch/$s_!CX1k!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CX1k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png" width="767" height="458" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:458,&quot;width&quot;:767,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CX1k!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png 424w, https://substackcdn.com/image/fetch/$s_!CX1k!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png 848w, https://substackcdn.com/image/fetch/$s_!CX1k!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png 1272w, https://substackcdn.com/image/fetch/$s_!CX1k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63aee96d-f0c2-4f2b-a0b4-c927342d9915_767x458.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure 2.1 &#8211; Architectural diagram of Skip-Gram model</p><p><strong>Step 1: Vocabulary and One-Hot Encoding</strong></p><ul><li><p>Define the vocabulary V, which is the set of unique words in the corpus.</p></li><li><p>Assign a unique index to each word in the vocabulary.</p></li><li><p>Represent the input and output words using one-hot encoding.</p></li><li><p>The one-hot encoded vector has a single 1 at the index corresponding to the word and 0s everywhere else.</p></li><li><p>The size of the one-hot encoded vector is equal to the vocabulary size |V|.</p></li></ul><p><strong>Example:</strong></p><ul><li><p>Let's say our vocabulary V consists of the following words: ["NLP", "has", "advanced", "rapidly", "with", "the", "rise", "of", "LLM", "models"].</p></li><li><p>The one-hot encoded vector for the word "NLP" would be: [1, 0, 0, 0, 0, 0, 0, 0, 0, 0], where the vector has size |V| = 10.</p></li></ul><p><strong>Step 2: Input Layer</strong></p><ul><li><p>The input to the Skip-Gram model is a one-hot encoded vector representing the target word.</p></li><li><p>The input layer has |V| neurons, each representing a word in the vocabulary.</p></li><li><p>The input layer is fully connected to the hidden layer.</p></li></ul><p><strong>Example:</strong></p><ul><li><p>If the target word is "LLM", the input to the Skip-Gram model would be a one-hot encoded vector of size |V| = 10, with a 1 at the index corresponding to "LLM" and 0s everywhere else.</p></li></ul><p><strong>Step 3: Embedding Matrix</strong></p><ul><li><p>The embedding matrix is a matrix of size |V| &#215; N, where N is the dimensionality of the word embeddings.</p></li><li><p>Each row of the embedding matrix represents the word embedding for a specific word in the vocabulary.</p></li><li><p>The embedding matrix is denoted as Matrix W in the diagram.</p></li><li><p>The initialization of the embedding matrix is an important consideration in the Skip-Gram model. The choice of initialization can impact the model's performance and convergence during training. There are several common approaches to initializing the embedding matrix such as: Random Initialization, Pre-trained Embeddings, Xavier Initialization, and He Initialization.</p></li></ul><p><strong>Example:</strong></p><ul><li><p>If we choose an embedding dimensionality of N = 5, the embedding matrix would have size |V| &#215; N = 10 &#215; 5.</p></li></ul><p><strong>Step 4: Hidden Layer</strong></p><ul><li><p>The hidden layer is obtained by multiplying the one-hot encoded input vector x of size |V| &#215; 1 with the embedding matrix W of size |V| &#215; N.</p></li><li><p>The resulting hidden layer h is a vector of size 1 &#215; N, representing the word embedding of the target word.</p></li></ul><p><strong>Example:</strong></p><ul><li><p>Given the one-hot encoded input vector x for the target word "LLM" and the embedding matrix W, the hidden layer h would be computed as: h = x &#215; W, resulting in a vector of size 1 &#215; N.</p></li></ul><p><strong>Step 5: Context Matrix</strong></p><ul><li><p>The context matrix is a matrix of size N &#215; |V|, denoted as Matrix W' in the diagram.</p></li><li><p>It maps the word embeddings from the hidden layer to the output layer.</p></li></ul><p><strong>Step 6: Output Layer</strong></p><ul><li><p>The output layer predicts the probabilities of each word in the vocabulary being a context word for the target word.</p></li><li><p>It is obtained by multiplying the hidden layer h of size 1 &#215; N with the context matrix W' of size N &#215; |V|.</p></li><li><p>The resulting output is a vector of size 1 &#215; |V|, representing the probabilities of each word being a context word.</p></li></ul><p><strong>Step 7: Output Softmax</strong></p><ul><li><p>The output vector undergoes a softmax activation function to convert the scores into probabilities.</p></li><li><p>The softmax function normalizes the output vector, ensuring that the probabilities sum up to 1.</p></li></ul><p><strong>Example:</strong></p><ul><li><p>For the target word "LLM", the output vector after the softmax function would represent the probabilities of each word in the vocabulary being a context word. The words with higher probabilities, such as "models", "advanced", and "NLP", would be considered more likely to appear in the context of "LLM".</p></li></ul><p>The Skip-Gram model is trained using techniques like stochastic gradient descent to optimize the embedding matrix W and the context matrix W'. The objective is to maximize the probability of predicting the correct context words given the target word.</p><p><strong>Code:</strong></p><pre><code>import numpy as np

from collections import defaultdict

class Word2Vec:

def __init__(self, corpus, window_size, embedding_size, learning_rate, epochs):

self.corpus = corpus

self.window_size = window_size

self.embedding_size = embedding_size

self.learning_rate = learning_rate

self.epochs = epochs

self.word_to_id, self.id_to_word = self.create_vocabulary()

self.word_counts = self.count_words()

self.embeddings = self.initialize_embeddings()

self.context_weights = self.initialize_context_weights()

def create_vocabulary(self):

words = set(self.corpus)

word_to_id = {word: i for i, word in enumerate(words)}

id_to_word = {i: word for i, word in enumerate(words)}

return word_to_id, id_to_word

def count_words(self):

word_counts = defaultdict(int)

for word in self.corpus:

word_counts[word] += 1

return word_counts

def initialize_embeddings(self):

vocab_size = len(self.word_to_id)

embeddings = np.random.uniform(-0.5, 0.5, (vocab_size, self.embedding_size))

return embeddings

def initialize_context_weights(self):

vocab_size = len(self.word_to_id)

context_weights = np.random.uniform(-0.5, 0.5, (self.embedding_size, vocab_size))

return context_weights

def softmax(self, x):

exp_x = np.exp(x - np.max(x))

return exp_x / np.sum(exp_x)

def train(self):

for epoch in range(self.epochs):

loss = 0

for i in range(len(self.corpus)):

center_word = self.corpus[i]

center_word_id = self.word_to_id[center_word]

context_words = self.get_context_words(i)

for context_word in context_words:

context_word_id = self.word_to_id[context_word]

# Forward pass

hidden_layer = self.embeddings[center_word_id]

output_layer = np.dot(hidden_layer, self.context_weights)

output_probs = self.softmax(output_layer)

# Calculate error and update weights

error = output_probs.copy()

error[context_word_id] -= 1

# Backpropagation

grad_context_weights = np.outer(hidden_layer, error)

grad_embeddings = np.dot(error, self.context_weights.T)

# Update weights

self.context_weights -= self.learning_rate * grad_context_weights

self.embeddings[center_word_id] -= self.learning_rate * grad_embeddings

loss += -np.log(output_probs[context_word_id])

print(f"Epoch: {epoch+1}, Loss: {loss}")

def get_context_words(self, i):

start = max(0, i - self.window_size)

end = min(len(self.corpus), i + self.window_size + 1)

context_words = self.corpus[start:i] + self.corpus[i+1:end]

return context_words

# Example usage

corpus = ["NLP", "has", "advanced", "rapidly", "with", "the", "rise", "of", "LLM", "models"]

window_size = 2

embedding_size = 50

learning_rate = 0.01

epochs = 10

word2vec = Word2Vec(corpus, window_size, embedding_size, learning_rate, epochs)

word2vec.train()

# Get the vector representations of the context words

center_word = "rise"

center_word_id = word2vec.word_to_id[center_word]

context_words = word2vec.get_context_words(word2vec.corpus.index(center_word))

context_vectors = []

for context_word in context_words:

context_word_id = word2vec.word_to_id[context_word]

context_vector = word2vec.context_weights[:, context_word_id]

context_vectors.append(context_vector)

print(f"Context words for '{center_word}': {context_words}")

print(f"Vector representations of the context words:")

for i, context_vector in enumerate(context_vectors):

print(f"{context_words[i]}: {context_vector}")</code></pre><h4><strong>Code Explanation</strong></h4><p>This code implements the Word2Vec model using the Skip-Gram architecture from scratch. Here's a breakdown of the code:</p><p>The Word2Vec class is defined with the following attributes:</p><ul><li><p>corpus: The input corpus of words.</p></li><li><p>window_size: The size of the context window.</p></li><li><p>embedding_size: The dimensionality of the word embeddings.</p></li><li><p>learning_rate: The learning rate for weight updates.</p></li><li><p>epochs: The number of training epochs.</p></li><li><p>The create_vocabulary method creates a mapping between words and their corresponding IDs and vice versa.</p></li><li><p>The count_words method counts the frequency of each word in the corpus.</p></li><li><p>The initialize_embeddings method initializes the word embeddings matrix with random values.</p></li><li><p>The initialize_context_weights method initializes the context weights matrix with random values.</p></li><li><p>The softmax method applies the softmax activation function to the output layer.</p></li><li><p>The train method trains the Word2Vec model using the Skip-Gram architecture. It iterates over each word in the corpus, identifies the context words within the specified window, and performs forward and backward propagation to update the embeddings and context weights.</p></li><li><p>The get_context_words method retrieves the context words for a given center word.</p></li></ul><p><strong>Word2Vec using pre-trained model</strong></p><p>One of the advantages of Word2Vec is the availability of pre-trained models that have already been trained on large datasets. These pre-trained models can be readily used for various NLP tasks without the need for training from scratch. One popular library that provides pre-trained Word2Vec models and easy-to-use functionality is Gensim.</p><p>Here's an explanation of how to use pre-trained Word2Vec models with Gensim, along with code examples:</p><ul><li><p><strong>Installation:</strong> To use Gensim, you first need to install it. You can install Gensim using pip:</p></li></ul><p><code>pip install gensim</code></p><p>Gensim is a popular open-source library for unsupervised topic modeling and NLP, with a focus on handling large text collections. It provides efficient and scalable implementations of various algorithms, including Word2Vec, for learning word embeddings from text data. Gensim is designed to be easy to use and integrate with other libraries in the Python ecosystem.</p><ul><li><p><strong>Loading a pre-trained Model:</strong> Gensim provides several pre-trained Word2Vec models that you can easily load. One commonly used model is the Google News model, which was trained on a large corpus of Google News articles. The Google News Word2Vec model consists of 300-dimensional word vectors for approximately 3 million words and phrases. Here's an example of how to load the Google News model:</p></li></ul><pre><code>from gensim.models import KeyedVectors

# Load the pre-trained Google News model

model = KeyedVectors.load_word2vec_format(' path/to/GoogleNews-vectors-negative300.bin.gz ', binary=True)</code></pre><ul><li><p><strong>Accessing Word Vectors:</strong> Once the pre-trained model is loaded, you can access the word vectors using the [] operator. Here's an example:</p></li></ul><pre><code># Get the vector representation of a word

vector = model['natural']

print(vector)</code></pre><p>This will retrieve the vector representation of the word &#8220;natural&#8221; from the pre-trained model.</p><ul><li><p><strong>Word Similarity:</strong> Word2Vec embeddings capture semantic similarity between words. You can use the similarity() method to calculate the cosine similarity between two words:</p></li></ul><pre><code># Calculate similarity between words

similarity = model.similarity(&#8216;NLP&#8217;, &#8216;LLM&#8217;)

print(similarity)</code></pre><ul><li><p><strong>Finding Similar Words: </strong>You can find the most similar words to a given word using the most_similar() method:</p></li></ul><pre><code># Find the most similar words

similar_words = model.most_similar('LLM', topn=5)

print(similar_words)</code></pre><ul><li><p><strong>Analogy Reasoning:</strong> Word2Vec embeddings can also be used for analogy reasoning, where you can find the word that completes an analogy. The most_similar() method can be used with positive and negative words:</p></li></ul><pre><code># Perform analogy reasoning

result = model.most_similar(positive=['Natural', 'Language'], negative=['Processing'], topn=1)

print(result)</code></pre><p>This will find the word that completes the analogy &#8220;king - man + woman = ?&#8221; based on the learned word embeddings.</p><div><hr></div><h3>FastText</h3><p>FastText is an open-source library developed by Facebook AI Research (FAIR) for efficient learning of word embeddings and text classification. It is an extension of the Word2Vec model, designed to address some of its limitations and improve upon its performance. FastText has gained popularity due to its ability to handle large-scale datasets efficiently and its effectiveness in capturing subword information.</p><h4><strong>The Need for FastText</strong></h4><p>Word2Vec has been a groundbreaking model for learning word embeddings, but it has certain limitations. One of the main drawbacks of Word2Vec is its inability to handle out-of-vocabulary (OOV) words. When Word2Vec encounters a word that was not present in the training data, it cannot provide a meaningful representation for that word. This is problematic when dealing with morphologically rich languages or specialized domains with rare or novel words.</p><p>FastText addresses this limitation by introducing the concept of subword embeddings. Instead of learning embeddings only for whole words, FastText breaks words into smaller subword units, such as character n-grams. By representing words as a combination of these subword embeddings, FastText can generate embeddings for OOV words by leveraging the information from their subword components.</p><p>FastText's ability to represent words as character n-grams and incorporate subword information sets it apart from other word embedding models like Word2Vec. Let's illustrate this with an example using the sentence we've been working with throughout the chapter:</p><p><em>"LLM models have transformed NLP systems and NLP capabilities."</em></p><p>In the Word2Vec model, each word in this sentence would be treated as a separate entity, and the model would learn embeddings for each word based on its context. However, FastText takes a different approach by breaking down each word into character n-grams.</p><p>Let's consider the word "transformed" and represent it using character trigrams (n = 3):</p><p><code>&lt;tra, ran, ans, nsf, sfo, for, orm, rme, med, ed&gt;</code></p><p>FastText would also include the entire word as a separate token:</p><p><code>&lt;tra, ran, ans, nsf, sfo, for, orm, rme, med, ed&gt; and &lt;transformed&gt;</code></p><p>By representing words as character n-grams, FastText captures the morphological and subword information within the words. This is particularly advantageous in several scenarios:</p><p><strong>Importance of FastText:</strong></p><ul><li><p><strong>Handling OOV Words: </strong>In the given example, let's say the word "transformational" was not present in the training data. Word2Vec would treat it as an unknown word and assign a random or default embedding. However, FastText can generate an embedding for "transformational" by leveraging the subword information it learned from similar words like "transformed" and "transformation".</p></li></ul><p>FastText would break down "transformational" into character n-grams: &lt;tra, ran, ans, nsf, sfo, for, orm, rma, mat, ati, tio, ion, ona, nal, al&gt; By utilizing the learned embeddings for these subword components, FastText can generate a meaningful representation for the OOV word "transformational".</p><ul><li><p><strong>Capturing Morphological Relationships:</strong> FastText's subword-based approach allows it to capture morphological relationships between words. In the example sentence, the words "transformed" and "capabilities" share common subword patterns:</p></li></ul><ul><li><p>&#8220;transformed&#8221;: <code>&lt;tra, ran, ans, nsf, sfo, for, orm, rme, med, ed&gt;</code></p></li><li><p>&#8220;capabilities&#8221;: <code>&lt;cap, apa, pab, abi, bil, ili, lit, iti, tie, ies, es&gt;</code></p></li></ul><p>By recognizing these shared subword components, FastText can understand the morphological similarity between these words and learn similar embeddings for them. This is particularly useful in languages with rich morphology, where words can have multiple inflected forms.</p><ul><li><p><strong>Efficient Representation and Computation:</strong> FastText's character n-gram representation allows for efficient storage and computation of word embeddings. Instead of learning separate embeddings for each word, FastText learns embeddings for the subword components. This reduces the number of parameters to be learned and allows for faster training and inference. In the example sentence, FastText would learn embeddings for the character n-grams and the entire words. During inference, the embedding for a word like "transformed" would be computed by combining the embeddings of its subword components, which is computationally efficient compared to storing and retrieving embeddings for every possible word.</p></li></ul><p><strong>Other important points about FastText are:</strong></p><ul><li><p><strong>Text Classification:</strong> In addition to learning word embeddings, FastText also provides a simple and efficient method for text classification. It represents text as a bag of word embeddings and uses a linear classifier to predict the class labels. FastText's text classification approach has shown competitive performance compared to more complex deep learning models, especially when dealing with large datasets and multiple classes.</p></li><li><p><strong>Multilingual Support:</strong> FastText has been trained on a wide range of languages and has pre-trained models available for many languages. This multilingual support allows users to easily apply FastText to various language-specific tasks and benefit from the learned embeddings across different languages.</p></li><li><p><strong>Integration and Extensibility:</strong> FastText is open-source and designed to be easily integrated into existing software systems. It provides a simple and intuitive API for training and using word embeddings. FastText can also be extended and customized to suit specific requirements, making it a flexible tool for researchers and practitioners.</p></li></ul><p><strong>Code:</strong></p><pre><code>from gensim.models import FastText

# Training data

sentences = [

"LLM models have transformed NLP systems and NLP capabilities.",

"The development of LLM models has revolutionized the field of NLP.",

"FastText is an efficient library for learning word embeddings.",

"FastText captures subword information and handles out-of-vocabulary words.",

]

# Train FastText model

model = FastText(sentences, vector_size=100, window=5, min_count=1, workers=4)

# Get the vocabulary

vocabulary = list(model.wv.key_to_index.keys())

print("Vocabulary:", vocabulary)

# Get the word vector for a word

word = "NLP"

vector = model.wv[word]

print(f"Vector for '{word}': {vector}")

# Find most similar words

similar_words = model.wv.most_similar(word, topn=5)

print(f"Most similar words to '{word}':")

for similar_word, similarity in similar_words:

print(f"- {similar_word}: {similarity}")

# Find similarity between words

word1 = "LLM"

word2 = "NLP"

similarity = model.wv.similarity(word1, word2)

print(f"Similarity between '{word1}' and '{word2}': {similarity}")

# Handle out-of-vocabulary words

oov_word = "transformational"

oov_vector = model.wv[oov_word]

print(f"Vector for OOV word '{oov_word}': {oov_vector}")</code></pre><p><strong>Description:</strong></p><p>Now, let's go through the detailed description to understand the code:</p><ul><li><p>We start by importing the FastText class from the gensim.models module. Gensim is a popular library for topic modeling and word embeddings.</p></li><li><p>We define our training data as a list of sentences. In this example, we have a small dataset related to LLM models and NLP.</p></li><li><p>We create an instance of the FastText model and train it on the provided sentences. The vector_size parameter specifies the dimensionality of the word vectors, window determines the context window size, min_count sets the minimum frequency threshold for words to be included in the vocabulary, and workers specifies the number of worker threads to use during training.</p></li><li><p>After training, we can access the vocabulary of the model using model.wv.key_to_index.keys(). We print the vocabulary to see the unique words in our dataset.</p></li><li><p>To get the word vector for a specific word, we use model.wv[word]. In this example, we retrieve the vector for the word "NLP" and print it.</p></li><li><p>We can find the most similar words to a given word using model.wv.most_similar(). We specify the word of interest and the number of top similar words to retrieve (topn). The method returns a list of tuples containing the similar words and their similarity scores. We print the most similar words to "NLP".</p></li><li><p>To calculate the similarity between two words, we use model.wv.similarity(). We provide the two words, "LLM" and "NLP", and print their cosine similarity score.</p></li><li><p>One of the advantages of FastText is its ability to handle out-of-vocabulary (OOV) words. We demonstrate this by trying to retrieve the vector for the word "transformational", which is not present in our training data. FastText can generate a vector for OOV words by leveraging the subword information learned during training.</p><div><hr></div></li></ul><h3>GloVe</h3><p>GloVe (Global Vectors for Word Representation) is an unsupervised learning algorithm for obtaining vector representations for words. It was introduced by Pennington et al. in 2014 as an improvement over existing word embedding techniques like Word2Vec and FastText. GloVe aims to capture both the global statistics of word co-occurrences in a corpus and the local context information.</p><h4><strong>The Need for GloVe</strong></h4><p>Previous word embedding methods, such as Word2Vec and FastText, have shown success in capturing semantic and syntactic relationships between words. However, these methods primarily rely on local context information, considering only the words within a fixed window size. They do not explicitly take into account the global co-occurrence statistics of words across the entire corpus.</p><p>GloVe addresses this limitation by incorporating both local and global information in the learning process. It leverages the global word-word co-occurrence matrix to capture the overall statistical patterns in the corpus. By doing so, GloVe aims to learn word vectors that better capture the semantic relationships between words.</p><p><strong>Importance of GloVe:</strong></p><ul><li><p><strong>Capturing Global Co-occurrence Statistics: </strong>GloVe leverages the global word-word co-occurrence matrix to capture overall semantic patterns and relationships between words, considering their broader context and associations.</p></li><li><p><strong>Efficient Training: </strong>GloVe's weighted least squares objective function enables efficient training of word vectors, even on large datasets, making it scalable for real-world applications.</p></li><li><p><strong>Improved Word Similarity and Analogy Performance: </strong>GloVe learns word vectors that effectively reflect semantic relationships, outperforming other methods on benchmark datasets for word similarity and analogy tasks.</p></li><li><p><strong>Interpretation of Vector Dimensions: </strong>GloVe word vectors exhibit interpretable dimensions that can capture specific semantic concepts, such as gender, sentiment, or part-of-speech information.</p></li><li><p><strong>Integration with Downstream Tasks: </strong>GloVe word vectors can be easily integrated into various downstream NLP tasks, serving as input features to enhance performance and generalization capabilities of machine learning models.</p></li></ul><p>Now that we have discussed the concept of GloVe and its importance in capturing both local and global word relationships, let's dive into the technical details of how GloVe learns word vectors using the co-occurrence information.</p><p>The core idea behind GloVe is to minimize the difference between the dot product of word vectors and the logarithm of their co-occurrence probabilities. This is achieved through the following objective function:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uwpb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uwpb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png 424w, https://substackcdn.com/image/fetch/$s_!uwpb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png 848w, https://substackcdn.com/image/fetch/$s_!uwpb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png 1272w, https://substackcdn.com/image/fetch/$s_!uwpb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uwpb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png" width="840" height="96" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:96,&quot;width&quot;:840,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:25184,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uwpb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png 424w, https://substackcdn.com/image/fetch/$s_!uwpb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png 848w, https://substackcdn.com/image/fetch/$s_!uwpb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png 1272w, https://substackcdn.com/image/fetch/$s_!uwpb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99b27162-0f71-4f54-8ac6-f12d8d185df0_840x96.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p></p><p>Where,</p><p><code>V </code>is the vocabulary size,</p><p><code>Xij</code> is the number of times word,</p><p><code>i </code>appears in the context of word <code>j,</code></p><p><code>fXij </code>is a weighting function,</p><p><code>wi </code>and <code>wj </code>are the word vectors, and</p><p><code>bi</code> and<code> bj</code> are the bias terms.</p><p>To illustrate the process of learning GloVe word vectors, let's walk through a step-by-step example using the sentence: "LLM models have transformed NLP systems and NLP capabilities." We'll use a context window size of 4 for this example.</p><p><strong>Step 1: Preprocess the text data</strong></p><ul><li><p>Tokenize the text into words</p></li><li><p>Remove punctuation and convert to lowercase</p></li><li><p>Create a vocabulary of unique words</p></li></ul><p><strong>Preprocessed:</strong><code> ["llm", "models", "have", "transformed", "nlp", "systems", "and", "nlp", "capabilities"]</code></p><p><strong>Vocabulary:</strong> <code>["llm", "models", "have", "transformed", "nlp", "systems", "and", "capabilities"]</code></p><p><strong>Step 2: Construct the co-occurrence matrix</strong></p><ul><li><p>Define a context window size (in this case, 4)</p></li><li><p>Iterate through the preprocessed text and count the co-occurrences of word pairs within the context window</p></li><li><p>Create a symmetric co-occurrence matrix X of size V * V, where V is the vocabulary size</p></li></ul><p>Co-occurrence matrix X (window size = 4):</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UyTE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UyTE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png 424w, https://substackcdn.com/image/fetch/$s_!UyTE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png 848w, https://substackcdn.com/image/fetch/$s_!UyTE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png 1272w, https://substackcdn.com/image/fetch/$s_!UyTE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UyTE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png" width="1320" height="738" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0595160e-8705-481d-9835-2c05eb944057_1320x738.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:738,&quot;width&quot;:1320,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:125142,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UyTE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png 424w, https://substackcdn.com/image/fetch/$s_!UyTE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png 848w, https://substackcdn.com/image/fetch/$s_!UyTE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png 1272w, https://substackcdn.com/image/fetch/$s_!UyTE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0595160e-8705-481d-9835-2c05eb944057_1320x738.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Step 3: Define the weighting function </strong>fXij</p><ul><li><p>Choose a weighting function to assign lower weights to rare and frequent co-occurrences</p></li><li><p>Common choice: fXij=min1, Xij / xmax</p></li></ul><blockquote><p>Where, xmax is a threshold and is a parameter (e.g., 0.75)</p></blockquote><p>Assuming xmax = 100 and = 0.75, the weighted co-occurrence matrix fXij remains the same as the co-occurrence matrix X in this example.</p><p><strong>Step 4: Initialize word vectors and biases</strong></p><ul><li><p>Randomly initialize word vectors wi and wj of dimension d (e.g., 100) for each word in the vocabulary</p></li><li><p>Initialize bias terms bi and bj to zero</p></li></ul><p><strong>Step 5: Train the GloVe model</strong></p><ul><li><p>Define the number of training epochs and learning rate</p></li><li><p>Iterate through each non-zero entry in the weighted co-occurrence matrix $f(X_{ij})$</p></li><li><p>For each entry (i, j):</p><ul><li><p>Compute the dot product of word vectors: wiT wj</p></li><li><p>Compute the error term: eij=wiT wj+ bi+ bj -log Xij</p></li><li><p>Update the word vectors and biases using gradient descent:</p><ul><li><p><code>wiwi- &#951;.eij.wj</code></p></li><li><p><code>wjwj- &#951;.eij.wi</code></p></li><li><p><code>bibi- &#951;.eij</code></p></li><li><p><code>bjbj- &#951;.eij</code></p></li></ul></li></ul></li><li><p>Repeat for the specified number of epochs</p></li></ul><p>Assuming a learning rate of 0.01 and 50 training epochs, the GloVe model will learn the word vectors wi and wj that capture the semantic relationships between words based on their co-occurrence patterns.</p><p><strong>Step 6: Evaluate the learned word vectors</strong></p><ul><li><p>Compute the cosine similarity between word vectors to find similar words</p></li><li><p>Perform word analogy tasks using vector arithmetic</p></li><li><p>Use the learned word vectors as features for downstream NLP tasks</p></li></ul><p>After training, the cosine similarity between the word vectors of "nlp" and "capabilities" might be high (e.g., 0.85), indicating their semantic relatedness. Word analogy tasks can also be performed, such as "nlp" - "systems" + "models" &#8776; "deep learning".</p><p>The learned GloVe word vectors can be used for various NLP applications, such as text classification, sentiment analysis, and information retrieval, as they capture meaningful semantic relationships between words based on their global co-occurrence statistics.</p><p><strong>Code:</strong></p><pre><code>import numpy as np

from sklearn.metrics.pairwise import cosine_similarity

def create_cooccurrence_matrix(corpus, vocab_size, window_size):

cooccurrence_matrix = np.zeros((vocab_size, vocab_size))

for i in range(len(corpus)):

for j in range(max(0, i - window_size), min(len(corpus), i + window_size + 1)):

if i != j:

cooccurrence_matrix[corpus[i], corpus[j]] += 1

return cooccurrence_matrix

def glove(cooccurrence_matrix, embedding_size, learning_rate, epochs):

vocab_size = cooccurrence_matrix.shape[0]

word_vectors = np.random.uniform(-0.5, 0.5, (vocab_size, embedding_size))

context_vectors = np.random.uniform(-0.5, 0.5, (vocab_size, embedding_size))

biases = np.random.uniform(-0.5, 0.5, (vocab_size,))

global_cooccurrences = np.sum(cooccurrence_matrix, axis=1)

for _ in range(epochs):

for i in range(vocab_size):

for j in range(vocab_size):

if cooccurrence_matrix[i, j] &gt; 0:

weight = (cooccurrence_matrix[i, j] / global_cooccurrences[i]) ** 0.75

error = np.dot(word_vectors[i], context_vectors[j]) + biases[i] + biases[j] - np.log(cooccurrence_matrix[i, j])

word_vectors[i] -= learning_rate * error * context_vectors[j] * weight

context_vectors[j] -= learning_rate * error * word_vectors[i] * weight

biases[i] -= learning_rate * error * weight

biases[j] -= learning_rate * error * weight

return word_vectors

# Example usage

corpus = [

"LLM models have transformed NLP systems and NLP capabilities.",

"NLP techniques are widely used in various applications.",

"Deep learning has revolutionized the field of NLP."

]

vocab = {}

preprocessed_corpus = []

for sentence in corpus:

words = sentence.lower().replace('.', '').split()

preprocessed_sentence = []

for word in words:

if word not in vocab:

vocab[word] = len(vocab)

preprocessed_sentence.append(vocab[word])

preprocessed_corpus.append(preprocessed_sentence)

vocab_size = len(vocab)

window_size = 4

embedding_size = 50

learning_rate = 0.01

epochs = 100

cooccurrence_matrix = create_cooccurrence_matrix(np.array([item for sublist in preprocessed_corpus for item in sublist]), vocab_size, window_size)

word_vectors = glove(cooccurrence_matrix, embedding_size, learning_rate, epochs)

word = "nlp"

word_index = vocab[word]

word_vector = word_vectors[word_index]

similarities = cosine_similarity([word_vector], word_vectors)[0]

most_similar = similarities.argsort()[-5:][::-1]

print(f"Words similar to '{word}':")

for index in most_similar:

for w, i in vocab.items():

if i == index:

print(f"- {w}: {similarities[index]}")</code></pre><p><strong>Description:</strong></p><ul><li><p>The create_cooccurrence_matrix function takes the corpus, vocabulary size, and window size as input and creates the co-occurrence matrix. It iterates over the corpus and counts the co-occurrences of words within the specified window size.</p></li><li><p>The glove function implements the GloVe algorithm. It takes the co-occurrence matrix, embedding size, learning rate, and number of epochs as input. The word vectors and context vectors are randomly initialized, and biases are also initialized.</p></li><li><p>Inside the training loop, the code iterates over each pair of words in the co-occurrence matrix. If the co-occurrence count is greater than zero, it calculates the weight using the local co-occurrence count divided by the global co-occurrence count of the target word, raised to the power of 0.75.</p></li><li><p>The error term is calculated by taking the dot product of the word vector and context vector, adding the biases, and subtracting the logarithm of the co-occurrence count.</p></li><li><p>The word vectors, context vectors, and biases are updated using gradient descent, multiplying the learning rate, error term, and weight.</p></li><li><p>After training, the word vectors capture the semantic relationships between words based on their co-occurrences.</p></li><li><p>The example usage section preprocesses the corpus by tokenizing the sentences, creating a vocabulary, and converting words to their corresponding indices.</p></li><li><p>The co-occurrence matrix is created using the create_cooccurrence_matrix function, and the GloVe model is trained using the glove function.</p></li><li><p>Finally, the code demonstrates how to find similar words by selecting a target word, retrieving its vector representation, and calculating the cosine similarity with all other word vectors. The top-5 most similar words are printed along with their similarity scores.</p><p></p><div><hr></div><h3>Embeddings (ELMo, BERT) and Contextual Representation</h3><p>The evolution of word embeddings from static representations, such as Word2Vec and GloVe, to contextual embeddings has been a significant milestone in natural language processing (NLP). Contextual embeddings like ELMo and BERT have redefined how models understand language by capturing context-sensitive meaning for words and phrases, even in ambiguous or complex contexts. This section explores the architecture, design, pre-training, and fine-tuning techniques for ELMo and BERT, providing a foundation for understanding their impact on NLP tasks.</p><div><hr></div><h4>Architecture and Design of ELMo and BERT</h4><p><strong>ELMo (Embeddings from Language Models)</strong></p><ul><li><p>ELMo generates word embeddings by incorporating context from the entire sentence, allowing the representation of words to change based on their usage.</p></li><li><p>It uses a bidirectional LSTM architecture, processing text both forwards and backwards, and stacks multiple layers for deeper contextual understanding.</p></li><li><p>ELMo embeddings are derived from all layers of the network, where lower layers capture syntax and higher layers capture semantic features.</p></li><li><p>The design focuses on pre-trained deep language models that are task-agnostic, making ELMo adaptable to various downstream applications by integrating into existing pipelines.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HAo7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HAo7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png 424w, https://substackcdn.com/image/fetch/$s_!HAo7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png 848w, https://substackcdn.com/image/fetch/$s_!HAo7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png 1272w, https://substackcdn.com/image/fetch/$s_!HAo7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HAo7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png" width="1426" height="745" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:745,&quot;width&quot;:1426,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;ELMo - Wikipedia&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="ELMo - Wikipedia" title="ELMo - Wikipedia" srcset="https://substackcdn.com/image/fetch/$s_!HAo7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png 424w, https://substackcdn.com/image/fetch/$s_!HAo7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png 848w, https://substackcdn.com/image/fetch/$s_!HAo7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png 1272w, https://substackcdn.com/image/fetch/$s_!HAo7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc099be01-bdb5-45bd-b44e-6c488534eb44_1426x745.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></li></ul><p><strong>BERT (Bidirectional Encoder Representations from Transformers)</strong></p><ul><li><p>BERT is based on the transformer architecture and leverages self-attention mechanisms to model the full context of words within a sentence.</p></li><li><p>Unlike traditional unidirectional models, BERT uses bidirectional attention, enabling it to consider both preceding and succeeding words simultaneously.</p></li><li><p>Its design includes multiple layers of transformers, with each layer refining representations using multi-headed self-attention and feedforward neural networks.</p></li><li><p>BERT&#8217;s flexibility is highlighted by its ability to handle sentence-pair tasks, such as question answering and textual entailment, in addition to single-sentence tasks.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oZK8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oZK8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png 424w, https://substackcdn.com/image/fetch/$s_!oZK8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png 848w, https://substackcdn.com/image/fetch/$s_!oZK8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png 1272w, https://substackcdn.com/image/fetch/$s_!oZK8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oZK8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png" width="1456" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c348862-f7ce-45db-902e-863d398536fe_1646x658.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Paper Walkthrough: Bidirectional Encoder Representations from Transformers ( BERT)&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Paper Walkthrough: Bidirectional Encoder Representations from Transformers ( BERT)" title="Paper Walkthrough: Bidirectional Encoder Representations from Transformers ( BERT)" srcset="https://substackcdn.com/image/fetch/$s_!oZK8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png 424w, https://substackcdn.com/image/fetch/$s_!oZK8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png 848w, https://substackcdn.com/image/fetch/$s_!oZK8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png 1272w, https://substackcdn.com/image/fetch/$s_!oZK8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c348862-f7ce-45db-902e-863d398536fe_1646x658.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></li></ul><div><hr></div><h4>Pre-Training Approaches for Contextual Embeddings</h4><p><strong>ELMo Pre-Training</strong></p><ul><li><p>ELMo is pre-trained on a large corpus using a language modeling objective. Specifically, it employs two separate objectives: forward prediction (next word prediction) and backward prediction (previous word prediction).</p></li><li><p>This dual-directional training allows ELMo to capture both past and future context in its embeddings.</p></li></ul><p><strong>Example: Generating ELMo Embeddings</strong></p><pre><code>from allennlp.commands.elmo import ElmoEmbedder

# Initialize ELMo Embedder
elmo = ElmoEmbedder()

# Input sentence
sentence = ["This", "is", "an", "example", "sentence", "."]

# Generate embeddings
embeddings = elmo.embed_sentence(sentence)

# Each word has three vectors (from different layers)
for i, word_embedding in enumerate(embeddings):
    print(f"Word {sentence[i]}: {word_embedding.shape}")  # Example: (3, 1024)
</code></pre><p><strong>BERT Pre-Training</strong></p><ul><li><p>BERT introduces two novel pre-training objectives:</p><ul><li><p><strong>Masked Language Modeling (MLM):</strong> Randomly masks words in a sentence and trains the model to predict the masked words using context from both directions.</p></li><li><p><strong>Next Sentence Prediction (NSP):</strong> Trains the model to predict whether a given sentence follows another in a sequence, improving its ability to handle tasks involving sentence pairs.</p></li></ul></li><li><p>These objectives allow BERT to develop a nuanced understanding of both sentence structure and inter-sentence relationships.</p><p></p><p><strong>Generating BERT Embeddings</strong></p><pre><code>from transformers import BertTokenizer, BertModel
import torch

# Load pre-trained BERT tokenizer and model
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertModel.from_pretrained('bert-base-uncased')

# Input text
text = "This is an example sentence."

# Tokenize input text
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)

# Generate embeddings
with torch.no_grad():
    outputs = model(**inputs)

# Extract last hidden states
last_hidden_states = outputs.last_hidden_state

print("Shape of embeddings:", last_hidden_states.shape)
# Example: torch.Size([1, 7, 768]) -&gt; (batch_size, sequence_length, hidden_size)
</code></pre></li></ul><div><hr></div><h4>Fine-Tuning Contextual Embeddings</h4><p><strong>Fine-Tuning ELMo</strong></p><ul><li><p>ELMo embeddings are typically used as additional input features to downstream models.</p></li><li><p>By freezing or slightly adjusting the weights of the pre-trained ELMo model, its embeddings can be fine-tuned to specific tasks like named entity recognition (NER), sentiment analysis, or question answering.</p></li></ul><p><strong>Fine-Tuning BERT</strong></p><ul><li><p>BERT is fine-tuned by adding a task-specific head (e.g., classification, regression, or sequence generation layers) on top of the pre-trained transformer layers.</p></li><li><p>During fine-tuning, the entire model is updated to optimize task-specific objectives, making BERT highly adaptable to diverse NLP applications.</p></li><li><p>Techniques like gradient clipping and learning rate warm-up are often employed to stabilize training and prevent catastrophic forgetting of pre-trained knowledge.</p><pre><code>from transformers import BertForSequenceClassification, AdamW
from transformers import Trainer, TrainingArguments

# Load pre-trained BERT model for classification
model = BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)

# Define tokenizer and data
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')

# Example dataset
texts = ["I love this!", "This is awful."]
labels = [1, 0]  # Positive = 1, Negative = 0

# Tokenize data
encodings = tokenizer(texts, truncation=True, padding=True, max_length=128, return_tensors="pt")

# Prepare dataset
class Dataset(torch.utils.data.Dataset):
    def __init__(self, encodings, labels):
        self.encodings = encodings
        self.labels = labels

    def __len__(self):
        return len(self.labels)

    def __getitem__(self, idx):
        item = {key: torch.tensor(val[idx]) for key, val in self.encodings.items()}
        item['labels'] = torch.tensor(self.labels[idx])
        return item

dataset = Dataset(encodings, labels)

# Training arguments
training_args = TrainingArguments(
    output_dir='./results',
    num_train_epochs=3,
    per_device_train_batch_size=8,
    evaluation_strategy="epoch",
    save_steps=10_000,
    save_total_limit=2,
    logging_dir='./logs',
)

# Define trainer
trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=dataset,
)

# Train model
trainer.train()</code></pre><div><hr></div></li></ul></li></ul><h1>Conclusion</h1><p>Chapter 2 provided a comprehensive foundation for understanding text representation techniques and contextual embeddings, which are essential for modern NLP systems. We began with classical approaches like Bag-of-Words and TF-IDF, progressing to advanced distributed representations such as Word2Vec, GloVe, and FastText. These methods emphasized capturing semantic meaning and efficient representation of text.</p><p>The chapter further explored embeddings like ELMo and BERT, which introduced the power of contextualized word representations. By discussing their architectures, pre-training strategies, and fine-tuning methodologies, we established a strong understanding of how these embeddings capture rich linguistic context and contribute to various NLP applications.</p><p>As the landscape of text representation has evolved, so too has the need for models capable of leveraging these representations to generate coherent and contextually relevant outputs. This brings us to the next chapter, where we focus on the progression from traditional statistical language models to modern neural network-based approaches.</p><div><hr></div><h1><strong>Connect with Me</strong></h1><ol><li><p>If you have any inquiries, feel free to reach out via message or email.</p></li></ol><blockquote><p><em><strong><a href="https://abonia1.github.io/">Website/Newletter</a></strong></em></p><p><em>Connect with me on<strong> <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p><p><em>Find me on<strong> <a href="https://github.com/Abonia1">Github</a></strong></em></p><p><em>Visit my technical channel on <strong><a href="https://www.youtube.com/@AboniaSojasingarayar">Youtube</a></strong></em></p></blockquote><h2></h2><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Open-Source Vision Language Models (VLMs) in Multimodal AI]]></title><description><![CDATA[Curated List - Open source - Multimodal AI]]></description><link>https://aboniasojasingarayar.substack.com/p/open-source-vision-language-models</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/open-source-vision-language-models</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Fri, 27 Dec 2024 08:01:06 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e8927271-040f-4759-99cf-eb59aac403c1_1118x826.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3><strong>Introduction to Multimodal AI and Vision Language Models (VLMs)</strong></h3><p>Multimodal AI refers to systems capable of processing and interpreting multiple forms of data&#8212;text, images, audio, and video. By combining these modalities, AI can develop a deeper understanding of complex information. <strong>Vision Language Models (VLMs)</strong> represent a subset of this field, focusing on processing visual data (images, videos) in conjunction with textual information.</p><p>In recent years, <strong>open-source VLMs</strong> have gained significant traction, challenging proprietary models with their performance and capabilities. This shift has democratized access to advanced multimodal AI, enabling researchers and developers to further explore and push the boundaries of AI.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!M08F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!M08F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png 424w, https://substackcdn.com/image/fetch/$s_!M08F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png 848w, https://substackcdn.com/image/fetch/$s_!M08F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png 1272w, https://substackcdn.com/image/fetch/$s_!M08F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!M08F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png" width="1118" height="826" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:826,&quot;width&quot;:1118,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:112793,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!M08F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png 424w, https://substackcdn.com/image/fetch/$s_!M08F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png 848w, https://substackcdn.com/image/fetch/$s_!M08F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png 1272w, https://substackcdn.com/image/fetch/$s_!M08F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc39409-96ec-4c78-a4b2-3941434d5209_1118x826.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Image: HuggingFace</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://aboniasojasingarayar.substack.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h3><strong>Key Players in Open-Source Vision Language Models</strong></h3><p>Several  <strong>open-source VLMs</strong> have emerged, each bringing unique capabilities to the table. Below ar</p><p>e some notable models:</p><h4><strong>Llama 3.2 Vision</strong></h4><p>An extension of Meta's Llama model, Llama 3.2 Vision is designed for image-text tasks, excelling in areas like captioning and image-based question answering.</p><p><strong>Model Architecture</strong><br>Llama 3.2-Vision builds upon the Llama 3.1 text-only model, an auto-regressive language model leveraging an optimized transformer architecture. The fine-tuned versions of the model utilize supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for safety and helpfulness. To handle image recognition tasks, Llama 3.2-Vision integrates a separately trained vision adapter with the pre-trained Llama 3.1 language model. This adapter includes a series of cross-attention layers that process image encoder representations and feed them into the core language model.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Hw_l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Hw_l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png 424w, https://substackcdn.com/image/fetch/$s_!Hw_l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png 848w, https://substackcdn.com/image/fetch/$s_!Hw_l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png 1272w, https://substackcdn.com/image/fetch/$s_!Hw_l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Hw_l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png" width="950" height="346" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:346,&quot;width&quot;:950,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:63111,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Hw_l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png 424w, https://substackcdn.com/image/fetch/$s_!Hw_l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png 848w, https://substackcdn.com/image/fetch/$s_!Hw_l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png 1272w, https://substackcdn.com/image/fetch/$s_!Hw_l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff189b4d-c54c-4b30-a39b-8459879171fa_950x346.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p>Parameters: 11B, 90B</p></li><li><p>Strengths: Visual content generation, chart, and diagram understanding</p></li><li><p>Limitations: Math-heavy tasks, language support restricted to English</p></li></ul><p><strong><a href="https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/">Learn more about Llama 3.2 Vision</a></strong></p><div><hr></div><h4><strong>NVLM 1.0</strong></h4><p>Developed by UC Berkeley, NVLM 1.0 is tailored for multimodal large language model tasks such as OCR and multimodal reasoning.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PQp6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PQp6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png 424w, https://substackcdn.com/image/fetch/$s_!PQp6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png 848w, https://substackcdn.com/image/fetch/$s_!PQp6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png 1272w, https://substackcdn.com/image/fetch/$s_!PQp6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PQp6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png" width="1124" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1124,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:325411,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PQp6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png 424w, https://substackcdn.com/image/fetch/$s_!PQp6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png 848w, https://substackcdn.com/image/fetch/$s_!PQp6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png 1272w, https://substackcdn.com/image/fetch/$s_!PQp6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F968ad748-f2b8-4419-80f2-4c8e609d6b87_1124x830.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure : NVLM-1.0 offers three architectural options: the cross-attention-based NVLM-X (top), the hybrid NVLM-H (middle), and the decoder-only NVLM-D (bottom). The dynamic high-resolution vision pathway is shared by all three models. However, different architectures process the image features from thumbnails and regular local tiles in distinct ways.</p><ul><li><p>Strengths: OCR, high-resolution image handling, multimodal reasoning</p></li><li><p>Special Features: NVLM-D (OCR), NVLM-H (hybrid reasoning)</p></li></ul><p><strong><a href="https://arxiv.org/abs/2409.11402">Learn more about NVLM 1.0</a></strong></p><div><hr></div><h4><strong>Molmo</strong></h4><p>Developed by the Allen Institute for AI, Molmo is a state-of-the-art VLM excelling across multiple benchmarks.It offers state-of-the-art performance across various benchmarks. An impressive model that can "point" to elements within images and perform complex multimodal tasks.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gI5C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gI5C!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png 424w, https://substackcdn.com/image/fetch/$s_!gI5C!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png 848w, https://substackcdn.com/image/fetch/$s_!gI5C!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!gI5C!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gI5C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png" width="1038" height="1120" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1120,&quot;width&quot;:1038,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:430360,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gI5C!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png 424w, https://substackcdn.com/image/fetch/$s_!gI5C!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png 848w, https://substackcdn.com/image/fetch/$s_!gI5C!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!gI5C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F16a514d3-6ae0-4fdb-85aa-e76b8b0faa87_1038x1120.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure :The <strong>Molmo</strong> architecture follows the simple and standard design of combining a language model with a vision encoder. Its strong performance is the result of a well-tuned training pipeline and our new <strong>PixMo</strong> data.</p><ul><li><p>Parameters: 1B, 7B, 72B</p></li><li><p>Strengths: Pointing in images, benchmark-leading performance</p></li><li><p>Limitations: Struggles with transparent images, requires preprocessing</p></li></ul><p><strong><a href="https://molmo.allenai.org/blog">Learn more about Molmo</a></strong></p><div><hr></div><h4><strong>Qwen2-VL</strong></h4><p>Qwen2-VL extends VLM capabilities to complex object relationships in scenes, with strong video analysis performance.</p><p><strong>Model Architecture</strong></p><p>A significant architectural enhancement in Qwen2-VL is the introduction of Naive Dynamic Resolution support. Unlike its predecessor, Qwen2-VL can handle images with arbitrary resolutions by mapping them to a dynamic number of visual tokens. This ensures alignment between the model's input and the intrinsic information in the images, providing a more human-like visual processing experience. As a result, the model can efficiently process images of varying clarity and size.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gdrC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gdrC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png 424w, https://substackcdn.com/image/fetch/$s_!gdrC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png 848w, https://substackcdn.com/image/fetch/$s_!gdrC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png 1272w, https://substackcdn.com/image/fetch/$s_!gdrC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gdrC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png" width="1456" height="946" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:946,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Qwen2-VL&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Qwen2-VL" title="Qwen2-VL" srcset="https://substackcdn.com/image/fetch/$s_!gdrC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png 424w, https://substackcdn.com/image/fetch/$s_!gdrC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png 848w, https://substackcdn.com/image/fetch/$s_!gdrC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png 1272w, https://substackcdn.com/image/fetch/$s_!gdrC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb38a6577-d1df-4ba3-a52d-2ed64a6a8350_1656x1076.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Multimodal Rotary Position Embedding (M-ROPE)</strong>: Decomposes positional embedding into parts to capture 1D textual, 2D visual, and 3D video positional information, enhancing its multimodal processing capabilities.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NRaD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NRaD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png 424w, https://substackcdn.com/image/fetch/$s_!NRaD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png 848w, https://substackcdn.com/image/fetch/$s_!NRaD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png 1272w, https://substackcdn.com/image/fetch/$s_!NRaD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NRaD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png" width="1456" height="292" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:292,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NRaD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png 424w, https://substackcdn.com/image/fetch/$s_!NRaD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png 848w, https://substackcdn.com/image/fetch/$s_!NRaD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png 1272w, https://substackcdn.com/image/fetch/$s_!NRaD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feeffb0c7-435b-48d2-8de6-79fd127a20e4_3059x613.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><ul><li><p>Strengths: Top-tier visual understanding, video comprehension</p></li><li><p>Limitations: Lacks built-in moderation, struggles with spatial reasoning</p></li></ul><p><strong><a href="https://github.com/QwenLM/Qwen2-VL">Learn more about Qwen2-VL</a></strong></p><div><hr></div><h4><strong>Pixtral</strong></h4><p>A 12-billion parameter model from Mistral, Pixtral is designed for both image and text processing, excelling in instruction-following tasks.</p><p><strong>Architecture</strong></p><p><strong>Variable Image Size:</strong> Pixtral is optimized for both speed and performance. It features a new vision encoder that natively supports variable image sizes:</p><ul><li><p>Images are passed through the vision encoder at their native resolution and aspect ratio, converting them into image tokens for each 16x16 patch.</p></li><li><p>These tokens are flattened into a sequence, with special [IMG BREAK] and [IMG END] tokens placed between rows and at the image&#8217;s end.</p></li><li><p>The [IMG BREAK] tokens help the model distinguish images with different aspect ratios, even when they have the same number of tokens.</p></li></ul><p>This approach enables Pixtral to effectively process complex high-resolution diagrams, charts, and documents while maintaining fast inference speeds for smaller images like icons, clipart, and equations.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2yun!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2yun!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png 424w, https://substackcdn.com/image/fetch/$s_!2yun!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png 848w, https://substackcdn.com/image/fetch/$s_!2yun!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png 1272w, https://substackcdn.com/image/fetch/$s_!2yun!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2yun!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png" width="1456" height="762" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:762,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Pixtral architecture&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Pixtral architecture" title="Pixtral architecture" srcset="https://substackcdn.com/image/fetch/$s_!2yun!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png 424w, https://substackcdn.com/image/fetch/$s_!2yun!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png 848w, https://substackcdn.com/image/fetch/$s_!2yun!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png 1272w, https://substackcdn.com/image/fetch/$s_!2yun!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba536207-2775-49fb-93c7-8c0bdae269c7_2738x1432.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Final Architecture:</strong><br>Pixtral is powered by a vision encoder trained from scratch to handle variable image sizes efficiently.</p><p>The final architecture consists of two main components:</p><ol><li><p><strong>Vision Encoder</strong> &#8211; This component tokenizes images.</p></li><li><p><strong>Multimodal Transformer Decoder</strong> &#8211; It predicts the next text token based on a sequence of text and images.</p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GlTC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GlTC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png 424w, https://substackcdn.com/image/fetch/$s_!GlTC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png 848w, https://substackcdn.com/image/fetch/$s_!GlTC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png 1272w, https://substackcdn.com/image/fetch/$s_!GlTC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GlTC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png" width="1456" height="799" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:799,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Detailed architecture&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Detailed architecture" title="Detailed architecture" srcset="https://substackcdn.com/image/fetch/$s_!GlTC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png 424w, https://substackcdn.com/image/fetch/$s_!GlTC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png 848w, https://substackcdn.com/image/fetch/$s_!GlTC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png 1272w, https://substackcdn.com/image/fetch/$s_!GlTC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3abe09f-4773-4829-a631-34f68db27a87_2138x1173.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Pixtral is trained to predict text tokens using interleaved image and text data, allowing the model to process any number of images of arbitrary sizes within its large context window of 128K tokens.</p><ul><li><p>Strengths: Multi-image processing, native resolution handling</p></li><li><p>Features: Context window of 128,000 tokens</p><p><strong><a href="https://arxiv.org/pdf/2410.07073">Learn more about Pixtral</a></strong></p><p></p><div><hr></div><h2>Comparison Chart</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WvF-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WvF-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png 424w, https://substackcdn.com/image/fetch/$s_!WvF-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png 848w, https://substackcdn.com/image/fetch/$s_!WvF-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png 1272w, https://substackcdn.com/image/fetch/$s_!WvF-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WvF-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png" width="948" height="710" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:710,&quot;width&quot;:948,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:143720,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WvF-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png 424w, https://substackcdn.com/image/fetch/$s_!WvF-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png 848w, https://substackcdn.com/image/fetch/$s_!WvF-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png 1272w, https://substackcdn.com/image/fetch/$s_!WvF-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F667845ce-6b1c-4e5a-88c7-88860b5cd763_948x710.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div></li></ul><h3><strong>Technical Aspects of Open-Source VLMs</strong></h3><p>Open-source VLMs leverage innovative techniques to enhance multimodal processing, including:</p><ol><li><p><strong>Dynamic Resolution Handling</strong>: Models like Qwen2-VL handle arbitrary image resolutions, offering a human-like processing experience by dynamically mapping images into visual tokens.</p></li><li><p><strong>Multimodal Rotary Position Embedding (M-ROPE)</strong>: A technique that improves the ability of models to understand 1D textual, 2D visual, and 3D video positional information.</p></li><li><p><strong>Tile-based Dynamic High-Resolution Image Processing</strong>: NVLM 1.0 boosts high-resolution image processing, particularly for multimodal reasoning and OCR tasks.</p></li><li><p><strong>Production-Grade Multimodality</strong>: NVLM 1.0 showcases enhanced text-only performance post-multimodal training, illustrating the benefits of multimodal data in LLMs.</p></li><li><p><strong>Dataset Quality and Diversity</strong>: Research suggests that dataset quality, rather than scale, significantly impacts pretraining success.</p><div><hr></div></li></ol><h3><strong>Challenges and Future Directions</strong></h3><p>While open-source VLMs have made great strides, challenges remain:</p><ol><li><p><strong>Math-Heavy Tasks</strong>: Many VLMs continue to struggle with tasks requiring high-level mathematical reasoning.</p></li><li><p><strong>Handling Transparent Images</strong>: Models like Molmo face difficulties when processing images with transparency, necessitating advanced preprocessing techniques.</p></li><li><p><strong>Ethical Considerations</strong>: Ensuring that VLMs have built-in moderation and safeguards is crucial, especially for models like Qwen2-VL, which currently lack such features.</p><div><hr></div></li></ol><h3><strong>Conclusion</strong></h3><p>The rise of powerful open-source VLMs marks a pivotal moment in the field of <strong>multimodal AI</strong>. By harnessing these models, we can push the boundaries of <strong>human-AI interaction</strong>,   solving complex, multimodal challenges. The future of multimodal AI is bright, and it will be exciting to see how these models evolve in real-world applications.</p><div><hr></div><h1><strong>Connect with Me</strong></h1><p>If you have any inquiries, feel free to reach out via message or email.</p><blockquote><p><em><strong><a href="https://abonia1.github.io/">Website/Newletter</a></strong></em></p><p><em>Connect with me on<strong> <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p><p><em>Find me on<strong> <a href="https://github.com/Abonia1">Github</a></strong></em></p><p><em>Visit my technical channel on <strong><a href="https://www.youtube.com/@AboniaSojasingarayar">Youtube</a></strong></em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Chapter 6 - Real-World Applications of RAG and LLMs]]></title><description><![CDATA[Hands-On real time RAG and LLM application]]></description><link>https://aboniasojasingarayar.substack.com/p/chapter-4-real-world-applications</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/chapter-4-real-world-applications</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Thu, 07 Nov 2024 07:01:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In the rapidly evolving landscape of artificial intelligence, the integration of Retrieval-Augmented Generation (RAG) and Large Language Models (LLMs) has opened up new horizons for real-world applications. This chapter delves into the practical use of RAG and LLMs across a wide array of domains, showcasing how these technologies are transforming industries and solving complex problems. By exploring case studies in various domains such as conversational AI, biomedical document understanding, and legal search, we aim to provide a comprehensive overview of how RAG and LLMs are being applied in real-world settings. Additionally, we will examine how these models are utilized to solve real-world problems, including open-domain question answering, long-form text generation, and multi-step reasoning. This exploration will not only highlight the potential of RAG and LLMs but also offer insights into the challenges and future directions of these technologies.</p><p>The chapter is structured to provide a detailed look at the applications of RAG and LLMs. It begins with a deep dive into case studies across different domains, providing you with a practical understanding of how these models are being deployed. This is followed by a section on solving real-world problems with RAG, offering examples of open-domain question answering, long-form text generation, and multi-step reasoning. To ensure a broad perspective, we also explore the landscape of vector-capable solutions, including approximate nearest neighbor libraries, vector databases, and cloud offerings. Finally, we delve into specific industries to see how RAG and LLMs are being utilized, providing a comprehensive understanding of their applications and impact.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this chapter, we will cover the following topics :</p><ul><li><p>Case studies in various domains:</p></li><li><p>Solving real-world problems with RAG</p></li><li><p>Landscape of vector-capable solutions</p></li><li><p>How RAG and LLMs are used in specific industries</p><div><hr></div><h1><strong>Find source code used in this chapter </strong></h1><blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://github.com/Abonia1/LLM-Engineering-Book&quot;,&quot;text&quot;:&quot;LLM Engineering Book&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://github.com/Abonia1/LLM-Engineering-Book"><span>LLM Engineering Book</span></a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m3Uh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m3Uh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png 424w, https://substackcdn.com/image/fetch/$s_!m3Uh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png 848w, https://substackcdn.com/image/fetch/$s_!m3Uh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png 1272w, https://substackcdn.com/image/fetch/$s_!m3Uh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m3Uh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png" width="1198" height="1414" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1414,&quot;width&quot;:1198,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:258100,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!m3Uh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png 424w, https://substackcdn.com/image/fetch/$s_!m3Uh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png 848w, https://substackcdn.com/image/fetch/$s_!m3Uh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png 1272w, https://substackcdn.com/image/fetch/$s_!m3Uh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F965dc4c3-7867-4ff4-b3cd-fb21dab8dc98_1198x1414.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div></blockquote><div><hr></div></li></ul><h1>Case studies in various domains</h1><p>In our earlier chapter, we explored the functionality of the RAG pipeline. If you haven't already, we highly suggest revisiting the preceding chapters before delving into this one, as it primarily focuses on the demonstration of implementation details. Below, we provide a simple illustration of how RAG operates in a question-answer format.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ag2_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ag2_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png 424w, https://substackcdn.com/image/fetch/$s_!Ag2_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png 848w, https://substackcdn.com/image/fetch/$s_!Ag2_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png 1272w, https://substackcdn.com/image/fetch/$s_!Ag2_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ag2_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png" width="986" height="538" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:538,&quot;width&quot;:986,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:66850,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ag2_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png 424w, https://substackcdn.com/image/fetch/$s_!Ag2_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png 848w, https://substackcdn.com/image/fetch/$s_!Ag2_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png 1272w, https://substackcdn.com/image/fetch/$s_!Ag2_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ecad586-2f74-49da-ba9a-4f7d4e63e507_986x538.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>NaiveRAG (Naive Retrieval Augmented Generation)</h4><p>This is the most basic form of RAG. It involves three steps:</p><ul><li><p>Retriever: Searches a large corpus of text (like Wikipedia) for passages relevant to a given prompt or question.</p></li><li><p>Generator: Uses the retrieved passages along with the prompt to generate a response using an LLM.</p></li><li><p>(Optional) Reranker: In some cases, a reranking step might be included where the retrieved passages are scored and re-ordered based on their relevance to the prompt.</p></li><li><p>NaiveRAG has limitations. The retrieved passages might not always be the most relevant, and the LLM might struggle to integrate them effectively.</p></li></ul><h4>AdvancedRAG</h4><p>This builds upon NaiveRAG by addressing its limitations. It can involve various techniques like:</p><p>Improved Retrieval Strategies: Optimizing how passages are searched and retrieved for better relevance.</p><p>Fine-tuning the Retriever: Adapting the retrieval process based on specific tasks or domains.</p><p>Advanced Prompt Engineering: Creating more specific prompts for the LLM to leverage the retrieved information effectively.</p><h4>ModularRAG</h4><p>This is the most flexible approach. It breaks down the RAG process into independent modules:</p><ul><li><p>Search Module: Responsible for retrieving relevant passages.</p></li><li><p>Memory Module: Manages the retrieved information.</p></li><li><p>Fusion Module: Combines the retrieved information with the prompt.</p></li><li><p>Routing Module: Decides which information to use based on the task.</p></li><li><p>Predict Module: Generates the final output using the LLM.</p></li><li><p>Task Adapter Module: Adapts the entire process for specific tasks.</p></li></ul><p>ModularRAG allows for more customization and control over each step in the reasoning process. Both NaiveRAG and AdvancedRAG can be seen as special cases of ModularRAG with a fixed set of modules.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0Kd6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0Kd6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png 424w, https://substackcdn.com/image/fetch/$s_!0Kd6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png 848w, https://substackcdn.com/image/fetch/$s_!0Kd6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png 1272w, https://substackcdn.com/image/fetch/$s_!0Kd6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0Kd6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png" width="800" height="455" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:455,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0Kd6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png 424w, https://substackcdn.com/image/fetch/$s_!0Kd6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png 848w, https://substackcdn.com/image/fetch/$s_!0Kd6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png 1272w, https://substackcdn.com/image/fetch/$s_!0Kd6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc34a5b10-1dc3-41f9-b0ca-07c366df506b_800x455.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Framework Supporting LLM powered Application</h2><p>Frameworks supporting RAG (Retrieval-Augmented Generation) and LLM (Large Language Models) application development include LangChain, LlamaIndex, Haystack, TinyLLM, Griptape,Embedchain and more.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!viB6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!viB6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png 424w, https://substackcdn.com/image/fetch/$s_!viB6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png 848w, https://substackcdn.com/image/fetch/$s_!viB6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png 1272w, https://substackcdn.com/image/fetch/$s_!viB6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!viB6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png" width="1022" height="710" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:710,&quot;width&quot;:1022,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:73972,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!viB6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png 424w, https://substackcdn.com/image/fetch/$s_!viB6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png 848w, https://substackcdn.com/image/fetch/$s_!viB6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png 1272w, https://substackcdn.com/image/fetch/$s_!viB6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c481d69-522c-447f-afad-2d714db5fdee_1022x710.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ul><li><p><strong>LangChain</strong> is an open-source framework designed to simplify the development of applications powered by large language models. It provides a comprehensive toolkit for building more complex and interactive LLM </p><p>applications, going beyond basic search and retrieval. LangChain's components include chains, which allow the chaining of components together, facilitating the use of PromptTemplates and LLMChains for interactive applications.</p></li><li><p><strong>LlamaIndex</strong> is a plug-and-play solution for search-centric applications, focusing on providing quick access to specific information within large datasets. It is more of a specialized framework compared to LangChain, which offers a broader range of applications requiring deeper customization.</p></li><li><p><strong>Haystack</strong> emerges as a comprehensive NLP framework, empowering developers to craft applications infused with cutting-edge NLP models and LLMs. With a diverse range of capabilities spanning question answering, answer generation, and semantic document search, Haystack heralds a new era of NLP-driven application development. Core concepts such as Pipelines and Nodes structure and process data, while Agents, powered by LLMs, navigate complex queries. Specialized tools augment agent capabilities, exemplified by calculators or WebRetrievers, while DocumentStores provide compatibility with various database technologies. Delve into the vast potential of NLP frameworks with Haystack's robust features and functionalities.</p><div><hr></div></li></ul><h2>Conversational AI</h2><p>Creating a conversational AI chatbot tailored to specific data needs involves several steps, from processing PDF documents to integrating a large language model (LLM) for generating responses. In this section we will walk you through the process, using HuggingFace Embeddings, FAISS for vector storage, and Ollama mistral model. The langchain library is instrumental in managing conversation chains, indexing data, and crafting prompt templates. By the end of this tutorial, you'll have built a RAG system with a conversational UI, capable of detecting hallucinations in the LLM's responses. Before we start we can install Ollma in our local machine for inference as follow:</p><p><strong>1.Download Ollama</strong></p><p>For linux users-To install Ollma on Linux, you can use the following command in your terminal. This command downloads and executes the installation script directly from the Ollma official site:</p><p><code>curl -fsSL https://ollama.com/install.sh | sh</code></p><p>For Windows users- Windows users can download Ollma by visiting the official Ollma website and following the link to the Windows download page:<a href="https://ollama.com/download/windows">https://ollama.com/download/windows</a></p><p>For mac users-https://ollama.com/download/mac</p><p><strong>2. Pulling the Model</strong></p><p>After installing Ollma, you can pull the model of your choice using the following command:</p><p><code>ollama pull mistral</code></p><p>This command downloads the specified model (in this case, "mistral") to your local machine, making it available for inference.</p><p>First, install necessary packages</p><p><code>!pip install pypdf langchain langchain-community tiktoken llama-cpp-python panel streamlit</code></p><p>Then, you need to process PDF documents to extract text and metadata. This step involves reading each page of the PDFs, extracting the text, and organizing it into a structured format for indexing.</p><pre><code>import PyPDF2

def prepare_docs(pdf_docs):

&nbsp;&nbsp;&nbsp;&nbsp;docs = []

&nbsp;&nbsp;&nbsp;&nbsp;metadata = []

&nbsp;&nbsp;&nbsp;&nbsp;content = []

&nbsp;&nbsp;&nbsp;&nbsp;for pdf in pdf_docs:

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;pdf_reader = PyPDF2.PdfReader(pdf)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;for index, page in enumerate(pdf_reader.pages):

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;doc_page = {'title': pdf + " page " + str(index + 1),

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;'content': page.extract_text()}

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;docs.append(doc_page)

&nbsp;&nbsp;&nbsp;&nbsp;for doc in docs:

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;content.append(doc["content"])

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;metadata.append({"title": doc["title"]})

&nbsp;&nbsp;&nbsp;&nbsp;print("Content and metadata are extracted from the documents")

&nbsp;&nbsp;&nbsp;&nbsp;return content, metadata</code></pre><p>Next, chunk the extracted content into smaller segments for easier processing and retrieval. This step uses the `RecursiveCharacterTextSplitter` from langchain to split the content based on a specified chunk size and overlap.</p><pre><code>from langchain.splitter import RecursiveCharacterTextSplitter

def get_text_chunks(content, metadata):

&nbsp;&nbsp;&nbsp;&nbsp;text_splitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;chunk_size=512,

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;chunk_overlap=256,

&nbsp;&nbsp;&nbsp;&nbsp;)

&nbsp;&nbsp;&nbsp;&nbsp;split_docs = text_splitter.create_documents(content, metadatas=metadata)

&nbsp;&nbsp;&nbsp;&nbsp;print(f"Documents are split into {len(split_docs)} passages")

&nbsp;&nbsp;&nbsp;&nbsp;return split_docs

Index the chunked documents into a FAISS-based vector database for efficient similarity search. This step uses HuggingFace Embeddings to generate embeddings for the documents, which are then stored in a FAISS database.

from langchain.vectorstore import HuggingFaceEmbeddings, FAISS

def ingest_into_vectordb(split_docs):

&nbsp;&nbsp;&nbsp;&nbsp;embeddings = HuggingFaceEmbeddings(model_name='sentence-transformers/all-MiniLM-L6-v2', model_kwargs={'device': 'cpu'})

&nbsp;&nbsp;&nbsp;&nbsp;db = FAISS.from_documents(split_docs, embeddings)

&nbsp;&nbsp;&nbsp;&nbsp;DB_FAISS_PATH = 'vectorstore/db_faiss'

&nbsp;&nbsp;&nbsp;&nbsp;db.save_local(DB_FAISS_PATH)

&nbsp;&nbsp;&nbsp;&nbsp;return db</code></pre><p>Configure a conversational chain for the Ollama model, integrating it with the vector database for information retrieval. This setup enhances the conversational experience by combining language generation with memory and retrieval functionalities.</p><p>To use another model replace the model name in Ollama(model="llama").</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Zq4C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Zq4C!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png 424w, https://substackcdn.com/image/fetch/$s_!Zq4C!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png 848w, https://substackcdn.com/image/fetch/$s_!Zq4C!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png 1272w, https://substackcdn.com/image/fetch/$s_!Zq4C!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Zq4C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png" width="1014" height="736" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:736,&quot;width&quot;:1014,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:180865,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Zq4C!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png 424w, https://substackcdn.com/image/fetch/$s_!Zq4C!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png 848w, https://substackcdn.com/image/fetch/$s_!Zq4C!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png 1272w, https://substackcdn.com/image/fetch/$s_!Zq4C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdf899af-6148-423e-94cd-d018ce5072d0_1014x736.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><pre><code>from langchain.memory import ConversationBufferMemory

from langchain.retrievalchain import ConversationalRetrievalChain

def get_conversation_chain(vectordb):

&nbsp;&nbsp;&nbsp;&nbsp;llm = Ollama(model="mistral")

&nbsp;&nbsp;&nbsp;&nbsp;retriever = vectordb.as_retriever()

&nbsp;&nbsp;&nbsp;&nbsp;memory = ConversationBufferMemory(

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;memory_key='chat_history', return_messages=True, output_key='answer'

&nbsp;&nbsp;&nbsp;&nbsp;)

&nbsp;&nbsp;&nbsp;&nbsp;conversation_chain = ConversationalRetrievalChain.from_llm(

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;llm=llm,

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;retriever=retriever,

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;memory=memory,

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return_source_documents=True

&nbsp;&nbsp;&nbsp;&nbsp;)

&nbsp;&nbsp;&nbsp;&nbsp;print("Conversational Chain created for the LLM using the vector store")

&nbsp;&nbsp;&nbsp;&nbsp;return conversation_chain</code></pre><p>Run below cell to prepare,chunk and vectorize the data.In below example I tested with my CV.Do not hesitate to use any pdf that you wants to work with.</p><pre><code>pdf_docs=["./data/CV.pdf"]

content, metadata = prepare_docs(pdf_docs)

split_docs = get_text_chunks(content, metadata)

vectordb=ingest_into_vectordb(split_docs)

Now , ask your Question.We created a conversational chain and now ready to chat with your own data.&nbsp;

### Question 1

user_question = "who is Abonia Sojasingarayar?"

response=conversation_chain({"question": user_question})

print("Q: ",user_question)

print("A: ",response['answer'])</code></pre><p>Output:</p><pre><code>Q:&nbsp; who is Abonia Sojasingarayar?
A: &nbsp; Abonia Sojasingarayar is a Machine Learning Scientist, Data Scientist, NLP Engineer, Computer Vision Engineer, AI Analyst, and Technical Writer. They have education from the Universit&#233; Pondicherry in India, IA School in Boulogne-Billancourt, France, and Institut F2I in Paris, France. Abonia has certifications from IBM and deeplearning.IA, and they are proficient in various tools and techniques related to their field such as Python, TensorFlow, GCP professional data engineer Badges, Watson Assistant, and RPA (Robotic Process Automation) among others. They have worked on projects involving API integration, machine learning pipeline development, and research engineering.

### Question 2&nbsp;

user_question = "where did she graduated?"

response=conversation_chain({"question": user_question})

print("Q: ",user_question)

print("A: ",response['answer'])

print("\nConversation Chain: \n",response)

Output:

Q:&nbsp; where did she graduated?

A: &nbsp; Abonia Sojasingarayar graduated from the Universit&#233; Pondicherry in India with a licence en technologie informatique et Ing&#233;nierie degree.</code></pre><p>Conversation Chain:&nbsp;</p><p><code>&nbsp;{'question': 'where did she graduated?', 'chat_history': [HumanMessage(content='who is Abonia Sojasingarayar?'), AIMessage(content=' Abonia Sojasingarayar is a Machine Learning Scientist, Data Scientist, NLP Engineer, Computer Vision Engineer, AI Analyst, and Technical Writer. They have education from the Universit&#233; Pondicherry in India&#8230;'), HumanMessage(content='where did she graduated?'), AIMessage(content=' Abonia Sojasingarayar graduated from the Universit&#233; Pondicherry in India with a licence en technologie informatique et Ing&#233;nierie degree.')], 'answer': ' Abonia Sojasingarayar graduated from the Universit&#233; Pondicherry in India with a licence en technologie informatique et Ing&#233;nierie degree.', 'source_documents': [Document(page_content='Abonia Sojasingarayar&nbsp; \n&nbsp; \n \n \nMachine Learning Scientist | Data Scientist | NLP Engineer | Computer Vision Engineer | AI \nAnalyst | Technical Writer&nbsp; \n &#8230;&#8230;", metadata={'title': './data/CV.pdf page 2'})]}</code></p><p>So, if we observe, when I query again without explicitly specifying the name, the model is now capable of recalling previous interactions, thanks to the integration of a memory buffer. This enhancement allows the model to retain information from past conversations, enabling it to provide more contextually relevant responses in subsequent interactions.</p><p>Finally, build a user interface for interacting with the chatbot. This UI allows users to ask questions related to their documents, with the application processing these questions, retrieving relevant information, and generating responses.</p><pre><code>import streamlit as st

def handle_userinput(user_question):

&nbsp;&nbsp;&nbsp;&nbsp;response = st.session_state.conversation({'question': user_question})

&nbsp;&nbsp;&nbsp;&nbsp;st.session_state.chat_history = response['chat_history']

&nbsp;&nbsp;&nbsp;&nbsp;for i, message in enumerate(st.session_state.chat_history):

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;if i % 2 == 0:

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;st.write(user_template.replace("{{MSG}}", message.content), unsafe_allow_html=True)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;else:

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;st.write(bot_template.replace("{{MSG}}", message.content), unsafe_allow_html=True)

def main():

&nbsp;&nbsp;&nbsp;&nbsp;st.set_page_config(page_title="Chat with your PDFs", page_icon=":books:")

&nbsp;&nbsp;&nbsp;&nbsp;st.header("Chat with multiple PDFs :books:")

&nbsp;&nbsp;&nbsp;&nbsp;user_question = st.text_input("Ask a question about your documents:")

&nbsp;&nbsp;&nbsp;&nbsp;if user_question:

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;handle_userinput(user_question)

&nbsp;&nbsp;&nbsp;&nbsp;with st.sidebar:

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;st.subheader("Your documents")

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;pdf_docs = st.file_uploader("Upload your PDFs here and click on 'Process'", accept_multiple_files=True)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;if st.button("Process"):

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;content, metadata = prepare_docs(pdf_docs)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;split_docs = get_text_chunks(content, metadata)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;vectorstore = ingest_into_vectordb(split_docs)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;st.session_state.conversation = get_conversation_chain(vectorstore)</code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m2dC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m2dC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png 424w, https://substackcdn.com/image/fetch/$s_!m2dC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png 848w, https://substackcdn.com/image/fetch/$s_!m2dC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png 1272w, https://substackcdn.com/image/fetch/$s_!m2dC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m2dC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png" width="996" height="602" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:602,&quot;width&quot;:996,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:52499,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!m2dC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png 424w, https://substackcdn.com/image/fetch/$s_!m2dC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png 848w, https://substackcdn.com/image/fetch/$s_!m2dC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png 1272w, https://substackcdn.com/image/fetch/$s_!m2dC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f7d6c1d-1fe2-4544-8478-6fe047991c71_996x602.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZxeZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png 424w, https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png 848w, https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png 1272w, https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png" width="1456" height="866" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:866,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png 424w, https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png 848w, https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png 1272w, https://substackcdn.com/image/fetch/$s_!ZxeZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbc9ed26-b9e2-4f95-bfa6-c51ee0e5e9be_1600x952.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Now, with our RAG conversational chatbot, we have the flexibility to upload any PDF document and immediately initiate a conversation.&nbsp;</p><div><hr></div><h2>Biomedical document understanding</h2><p>Biomedical document understanding through Large Language Models (LLMs) involves leveraging advanced AI technologies to process and comprehend vast amounts of medical text, such as research papers, clinical studies, and patient records. This capability is crucial for healthcare providers to stay updated with the latest medical information and make informed decisions. LLMs, specifically designed for healthcare applications, can analyze complex medical texts, extract meaningful information, and generate insights for healthcare professionals. They are distinguished by the databases they were trained on, with clinical LLMs focusing on medical literature for diagnostic support and biomedical LLMs facilitating fast and accessible biomedical text mining.LLMs in healthcare have numerous use cases, including processing extensive databases of medical literature and patient data, learning from historical cases, and providing insights for accurate and timely diagnosis. For example, a LLM can analyze a patient&#8217;s symptoms, medical history, and clinical findings to generate a personalized treatment plan, incorporating the latest research findings and treatment guidelines. This enhances patient care and outcomes by enabling healthcare professionals to make more informed decisions.</p><p>Pretrained models like BioBERT, ClinicalBERT, BlueBERT, and BioGPT have shown significant advancements in applying AI in the medical field. BioBERT, trained on large-scale biomedical corpora, excels in understanding complex medical texts and terminology, making it effective for tasks like disease prediction and drug-drug interaction analysis. ClinicalBERT, adapted from BioBERT, is fine-tuned on clinical notes for more accurate patient data analysis and decision support. BlueBERT offers a balanced understanding of both biomedical and clinical texts, making it versatile for various applications. BioGPT, a generative pretrained transformer model, is useful for generating coherent medical text.</p><p>Med-PaLM, a large-scale generalist biomedical AI system, stands out as a multimodal generative model designed to handle various types of biomedical data, including clinical language, medical imaging, and genomics. It leverages advances in language and multimodal foundation models, allowing for rapid adaptation to different tasks and settings. Med-PaLM achieves remarkable performance on a wide range of tasks within the MultiMedBench benchmark, often surpassing state-of-the-art specialist models. This model demonstrates promising potential for downstream data-scarce biomedical applications and has the ability to process inputs with multiple images during inference, effectively handling complex medical scenarios.</p><div><hr></div><h2>Legal Summarizer</h2><p>A Legal Summarizer that simplifies complex legal documents, making them easier to understand for various audiences, including lawyers and non-lawyers. It extracts and condenses the key points, legal terms, and essential information from legal documents, such as court cases, legal briefs, and legislation, into concise summaries. This is particularly useful for:</p><p>Lawyers: They can quickly grasp the core arguments, legal issues, and outcomes without sifting through lengthy documents. It helps in preparing for legal cases, advising clients, and researching legal topics more efficiently.</p><p>Non-lawyers: For individuals who need to understand legal documents for personal, business, or educational reasons but lack expertise in legal terminology. It provides a clear, accessible summary of legal documents, enabling them to make informed decisions or understand legal implications more effectively.</p><p>Legal Education and Research: Students and researchers can use Legal Summarizers to grasp complex legal concepts and cases without the need to read through entire volumes of legal texts. This aids in studying law, conducting legal research, and preparing for exams.</p><p>Legal Assistants and Support Staff: They can use summaries to provide clients or colleagues with essential information from legal documents, making it easier to convey complex legal issues in a straightforward manner.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!66rb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!66rb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png 424w, https://substackcdn.com/image/fetch/$s_!66rb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png 848w, https://substackcdn.com/image/fetch/$s_!66rb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png 1272w, https://substackcdn.com/image/fetch/$s_!66rb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!66rb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png" width="1456" height="649" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:649,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!66rb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png 424w, https://substackcdn.com/image/fetch/$s_!66rb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png 848w, https://substackcdn.com/image/fetch/$s_!66rb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png 1272w, https://substackcdn.com/image/fetch/$s_!66rb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd47d5e86-d0f8-4f41-9f56-786a49d82fdb_1600x713.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>To perform document summarization using LLMs with the LangChain library, you have three main options: Stuff, Map-Reduce, and Refine. In the coming section we will see the hands on guide for each method.</p><h3>Option 1: Stuff</h3><p>This method involves stuffing all your documents into a single prompt and passing it to an LLM.</p><p>Import necessary modules and define the prompt template.</p><p>Create an LLM chain with the defined prompt.</p><p>Define a StuffDocumentsChain that takes the LLM chain and combines all documents into a single prompt.</p><p>Run the summarization.</p><pre><code>from langchain.chains.combine_documents.stuff import StuffDocumentsChain

from langchain.chains.llm import LLMChain

from langchain.prompts import PromptTemplate

loader = PyPDFLoader("./data/Raptor-Agreement.pdf")

# Define prompt

prompt_template = """Write a concise summary of the following:

"{text}"

CONCISE SUMMARY:"""

prompt = PromptTemplate.from_template(prompt_template)

# Define LLM chain

llm = ChatOpenAI(temperature=0, model_name="gpt-3.5-turbo-16k")

llm_chain = LLMChain(llm=llm, prompt=prompt)

# Define StuffDocumentsChain

stuff_chain = StuffDocumentsChain(llm_chain=llm_chain, document_variable_name="text")

docs = loader.load()

print(stuff_chain.run(docs))</code></pre><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yTUn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yTUn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png 424w, https://substackcdn.com/image/fetch/$s_!yTUn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png 848w, https://substackcdn.com/image/fetch/$s_!yTUn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png 1272w, https://substackcdn.com/image/fetch/$s_!yTUn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yTUn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png" width="1456" height="196" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:196,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yTUn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png 424w, https://substackcdn.com/image/fetch/$s_!yTUn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png 848w, https://substackcdn.com/image/fetch/$s_!yTUn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png 1272w, https://substackcdn.com/image/fetch/$s_!yTUn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bbb9e3f-3105-4a20-a31a-2bc85dc7dd86_1600x215.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Due to the limitations of summarizing lengthy content, we may not be able to provide the full summary results here. However, you're welcome to clone the GitHub repository associated with this project. Within the repository, you'll find notebooks corresponding to each chapter. By running these notebooks, you'll be able to explore the complete results and gain a deeper understanding of the concepts discussed. Feel free to experiment.</p><h3>Option 2: Map-Reduce</h3><p>This method involves summarizing each document individually (map) and then combining these summaries into a final summary (reduce).</p><p><em>-Define the map and reduce prompts.</em></p><p><em>-Create an LLM chain for mapping each document to an individual summary.</em></p><p><em>-Use a ReduceDocumentsChain to combine the summaries.</em></p><p><em>-Optionally, use a MapReduceDocumentsChain to automate the process.</em></p><pre><code>from langchain.chains import MapReduceDocumentsChain, ReduceDocumentsChain

from langchain_text_splitters import CharacterTextSplitter

llm = ChatOpenAI(temperature=0)

# Map

map_template = """The following is a set of documents

{docs}

Based on this list of docs, please identify the main themes&nbsp;

Helpful Answer:"""

map_prompt = PromptTemplate.from_template(map_template)

map_chain = LLMChain(llm=llm, prompt=map_prompt)

# Reduce

reduce_template = """The following is set of summaries:

{docs}

Take these and distill it into a final, consolidated summary of the main themes.&nbsp;

Helpful Answer:"""

reduce_prompt = PromptTemplate.from_template(reduce_template)

# Run chain

reduce_chain = LLMChain(llm=llm, prompt=reduce_prompt)

combine_documents_chain = StuffDocumentsChain(

&nbsp;&nbsp;&nbsp;&nbsp;llm_chain=reduce_chain, document_variable_name="docs"

)

reduce_documents_chain = ReduceDocumentsChain(

&nbsp;&nbsp;&nbsp;&nbsp;combine_documents_chain=combine_documents_chain,

&nbsp;&nbsp;&nbsp;&nbsp;collapse_documents_chain=combine_documents_chain,

&nbsp;&nbsp;&nbsp;&nbsp;token_max=4000,

)

map_reduce_chain = MapReduceDocumentsChain(

&nbsp;&nbsp;&nbsp;&nbsp;llm_chain=map_chain,

&nbsp;&nbsp;&nbsp;&nbsp;reduce_documents_chain=reduce_documents_chain,

&nbsp;&nbsp;&nbsp;&nbsp;document_variable_name="docs",

&nbsp;&nbsp;&nbsp;&nbsp;return_intermediate_steps=False,

)

text_splitter = CharacterTextSplitter.from_tiktoken_encoder(

&nbsp;&nbsp;&nbsp;&nbsp;chunk_size=1000, chunk_overlap=0

)

split_docs = text_splitter.split_documents(docs)

print(map_reduce_chain.run(split_docs))</code></pre><p>Here is the output:<br><em>This document outlines the terms and conditions of an Agreement between Raptor Technologies, LLC (Raptor) and a Subscriber organization for access to Raptor's Subscription Services. The key themes include:</em></p><p><em>1. License and Terms: Raptor grants a limited, non-exclusive license to the Subscriber to use its Subscription Services subject to certain terms and conditions. The Subscriber is responsible for providing their own Internet access and equipment to use the Subscription Services.</em></p><p><em>2. Confidentiality: The Subscriber agrees to keep confidential any information related to the Subscription Services and Equipment provided by Raptor, except as expressly permitted.</em></p><p><em>3. Data Collection and Distribution: The Subscriber is prohibited from disclosing or making public individual's personally identifying information obtained through the Subscription Services except as required in the ordinary course of business or by applicable law.</em></p><p><em>4. Fees and Term: The Agreement has an initial term of one year, during which the Subscriber must pay the Annual Software Access Fee for each Campus that will utilize the Subscription Services. Upon termination, all amounts due to Raptor remain payable and all licenses granted under the Agreement terminate at the end of the pre-paid annual term.</em></p><p><em>5. Termination: The Subscriber may terminate the Agreement with 60 days' written notice prior to the end of the then-current term. Sections 1, 2, 3, 6, and 7 survive termination.</em></p><p><em>6. Disclaimers: Raptor does not guarantee or warrant any information made available within the Subscription Services, including determinations of an individual's registered sex offender status or custom alert status. The Subscriber is responsible for ensuring compliance with applicable laws and regulations related to data collection and distribution.</em></p><p><em>7. Other Provisions: The Agreement includes provisions related to acts beyond Raptor's control, lack of creation of partnership or agency relationship, and non-assignment by the Subscriber without consent. Contact information for written notices and effective date are also included.</em></p><h3>Option 3: Refine</h3><p>This method involves iteratively refining a summary based on new context.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CrEi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CrEi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png 424w, https://substackcdn.com/image/fetch/$s_!CrEi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png 848w, https://substackcdn.com/image/fetch/$s_!CrEi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png 1272w, https://substackcdn.com/image/fetch/$s_!CrEi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CrEi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png" width="1456" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CrEi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png 424w, https://substackcdn.com/image/fetch/$s_!CrEi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png 848w, https://substackcdn.com/image/fetch/$s_!CrEi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png 1272w, https://substackcdn.com/image/fetch/$s_!CrEi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6298c911-31cd-4757-a1f0-611533b4b181_1600x692.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>-Define the prompt template for refining.</em></p><p><em>-Load the summarize chain with the refine chain type.</em></p><p><em>-Run the summarization with the input documents.</em></p><pre><code>from langchain import load_summarize_chain, PromptTemplate

prompt_template = """Write a concise summary of the following:

"{text}"

CONCISE SUMMARY:"""

prompt = PromptTemplate.from_template(prompt_template)

chain = load_summarize_chain(llm, chain_type="refine")

chain.run(split_docs)</code></pre><p>Output:</p><p><em>" This document outlines the terms of a subscription agreement between Subscriber (district/school or organization) and Raptor Technologies LLC (Raptor) for access to Raptor's Subscription Services. The agreement grants Subscriber a limited, non-exclusive license to use the services in accordance with the agreement and applicable laws. Confidential information provided by Raptor must be kept confidential and not disclosed to third parties without prior written consent. Individual's personally identifying information obtained through the services must not be disclosed except as required by law or in the ordinary course of business. Subscriber is responsible for providing its own Internet access and equipment to use the services, and fees are payable annually in advance. The agreement has an initial term of one year, with automatic renewal unless written notice of non-renewal is given.\n\nRaptor disclaims all responsibility for determinations of an individual&#8217;s registered sex offender status or custom alert status based on the information conveyed in connection with the Subscription Services. Subscriber is solely responsible for such determinations and understands that information provided by Raptor is not intended to substitute for the determinations made by Subscriber and its employees and contractors.\n\nThe agreement may be amended only pursuant to a written agreement between the Parties. All terms and conditions of this Agreement shall be binding upon, inure to the benefit of, and be enforceable by, the Parties and their respective successors and permitted assigns. Raptor will not be in default of this Agreement for any performance failure caused by occurrences beyond Raptor&#8217;s reasonable control (including, but not limited to, acts of God). This Agreement does not create any right enforceable by any person not a party. Nothing in this Agreement shall create the relationship of partners or principal-agent between the parties. Subscriber may not assign this Agreement without the prior written consent of Raptor. The waiver or failure of Raptor to exercise in any respect any right provided for under this Agreement shall not be deemed a waiver of any further right under this Agreement."</em></p><p>Each of these methods has its use cases depending on the specific requirements of your document summarization task. The Stuff method is simpler but may not capture all nuances. Map-Reduce provides a more detailed approach but requires more setup. Refine offers a way to iteratively improve a summary based on additional context.</p><div><hr></div><h1>Solving real-world problems with RAG</h1><p>Retrieval-Augmented Generation (RAG) approach addresses the challenge of generating factually accurate and coherent long-form text for open-domain question answering (QA), long-form text generation, and multi-step reasoning tasks. By retrieving relevant documents and incorporating their information into the model's responses, RAG improves the relevance, factual correctness, and attribution of the generated text, making it particularly useful for domains requiring external knowledge such as science, medicine, and technical support.</p><h2>1.Open-domain QA</h2><p>Let's dive directly to the demo.LangChain provides multiple built-in document loaders, that work with PDF files, JSON files, or a Python file in your file directory.&nbsp; We can use LangChain&#8217;s PyPDFLoader to import your PDF seamlessly. Here we will be using the data directly from the website to ask questionsHere we will be using medium articles.</p><p>Install necessary packages</p><pre><code>!pip install nest_asyncio langchain_community langchain playwright html2text sentence-transformers faiss-cpu

from langchain_community.document_transformers import Html2TextTransformer

from langchain.text_splitter import CharacterTextSplitter

from langchain_community.embeddings import HuggingFaceEmbeddings

from langchain_community.vectorstores import FAISS

import nest_asyncio

nest_asyncio.apply()

# Articles to index

articles = ["https://medium.com/@abonia/bertscore-explained-in-5-minutes-0b98553bfb71",

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"https://medium.com/@abonia/document-based-llm-powered-chatbot-bb316009de93/",]

# Scrapes the blogs above

loader = AsyncChromiumLoader(articles)

docs = loader.load()

</code></pre><p>When our document is long, it&#8217;s necessary to split up our document text into chunks. There are various ways to split your text. Let&#8217;s just use the simplest method CharacterTextSplitter to split based on characters and measure chunk length by the number of characters.&nbsp;</p><pre><code># Converts HTML to plain text&nbsp;

html2text = Html2TextTransformer()

docs_transformed = html2text.transform_documents(docs)

# Chunk text

text_splitter = CharacterTextSplitter(chunk_size=100,&nbsp;

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;chunk_overlap=0)

chunked_documents = text_splitter.split_documents(docs_transformed)

# Load chunked documents into the FAISS index

db = FAISS.from_documents(chunked_documents,&nbsp;

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;HuggingFaceEmbeddings(model_name='sentence-transformers/all-mpnet-base-v2'))

retriever = db.as_retriever()</code></pre><p>The text chunks are then translated into numerical vectors through embeddings, allowing us to work with text data like semantic search in a computationally efficient manner. We can choose an embedding model provider like OpenAI, HuggingFaceEmbedding, Jina etc for this task.We then need to store our embedding vectors in a vector store, which allows us to search and retrieve the relevant vectors at query time.&nbsp;</p><p>We can expose the vector store in a retriever interface. To retrieve text, we can choose a search type like &#8220;similarity&#8221; to use similarity search in the retriever object where it selects text chunk vectors that are most similar to the question vector. k=2 lets us find the top 2 most relevant text chunk vectors.&nbsp; A RetrievalQA chain chains a large language model with our retriever interface. You can also define the chain type as one of the four options: &#8220;stuff,&#8221; &#8220;map reduce,&#8221; &#8220;refine,&#8221; &#8220;map_rerank.&#8221; The default chain_type=&#8221;stuff&#8221; incorporates ALL text from the documents into the prompt. The &#8220;map_reduce&#8221; type breaks texts into groups, poses the question to the LLM for each batch separately, and derives the ultimate answer based on the replies from each batch.&nbsp;</p><p>The &#8220;refine&#8221; type partitions texts into batches, presents the first batch to the LLM, and then submits the answer along with the second batch to the LLM. It progressively refines the answer by processing through all the batches. The &#8220;map-rerank&#8221; type divides texts into batches, submits each one to the LLM, returns a score indicating how comprehensively it answers the question, and determines the final answer based on the highest-scoring replies from each batch.</p><pre><code>from langchain import PromptTemplate

from langchain_community.llms import Ollama

from langchain.chains import RetrievalQA

prompt_template = """

### [INST] Instruction: Answer the question based on the medium article knowledge. Here is context to help:

{context}

### QUESTION:

{question} [/INST]

&nbsp;"""

# Create prompt from prompt template

prompt = PromptTemplate(

&nbsp;&nbsp;&nbsp;&nbsp;input_variables=["context", "question"],

&nbsp;&nbsp;&nbsp;&nbsp;template=prompt_template,

)

llm = Ollama(model="mistral")

qa = RetrievalQA.from_chain_type(llm=llm, chain_type="stuff", retriever=retriever, return_source_documents=True, chain_type_kwargs={"prompt": prompt_template},)

answer = qa.invoke("What is cosine similarity?")
</code></pre><p>Output</p><p><em>{'query': 'What is cosine similarity?',</em></p><p><em>&nbsp;'result': ' Cosine similarity is a measure of similarity between two non-zero vectors of an inner product space. It is computed as the cosine of the angle between them, which indicates how similar they are in direction. The result ranges from -1 to 1, with 1 indicating perfect similarity and 0 indicating orthogonal (perpendicular) vectors.',</em></p><p><em>&nbsp;'source_documents': [Document(page_content='The formula for cosine similarity is:\n\n&gt; similarity(A, B) = (A . B) / (||A|| ||B||)', metadata={'source': 'https://medium.com/@abonia/document-based-llm-powered-chatbot-bb316009de93/'}),</em></p><p><em>&nbsp;&nbsp;Document(page_content='Cosine similarity &#8212; This method measures the cosine of the angle between two\nvectors, which indicates how similar they are in direction. Cosine similarity\nranges from -1 to 1, with 1 indicating perfect similarity.', metadata={'source': 'https://medium.com/@abonia/document-based-llm-powered-chatbot-bb316009de93/'}),</em></p><p><em>&nbsp;&nbsp;Document(page_content='Cosine Similarity and Cosine Distance &#8212; credit', metadata={'source': 'https://medium.com/@abonia/document-based-llm-powered-chatbot-bb316009de93/'}),</em></p><p><em>&nbsp;&nbsp;Document(page_content='Where A and B are the two vectors being compared, . is the dot product of the\nvectors, and || || represents the Euclidean norm (magnitude) of the vectors.', metadata={'source': 'https://medium.com/@abonia/document-based-llm-powered-chatbot-bb316009de93/'})]}</em></p><p>Sample code to build RAG Chain</p><pre><code>llm_chain = LLMChain(llm=mistral_llm, prompt=prompt)

rag_chain = (&nbsp;

&nbsp;{"context": retriever, "question": RunnablePassthrough()}

&nbsp;&nbsp;&nbsp;&nbsp;| llm_chain

)

result = rag_chain.invoke("What is cosine similarity??")

print(result['text'])</code></pre><h2>2.Long-form text generation</h2><h2>Multi-step reasoning and Augmentation</h2><p>To implement multi-step reasoning in a Retrieval-Augmented Generation (RAG) application, we can follow a multi-stage retrieval process that combines different retrieval methods for improved overall quality. Augmentation involves the process of effectively integrating context from retrieved passages with the current generation task. Before discussing more on the augmentation process, augmentation stages, and augmentation data, here is a taxonomy of RAG's core components:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!my2M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!my2M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png 424w, https://substackcdn.com/image/fetch/$s_!my2M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png 848w, https://substackcdn.com/image/fetch/$s_!my2M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png 1272w, https://substackcdn.com/image/fetch/$s_!my2M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!my2M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png" width="1456" height="613" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:613,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!my2M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png 424w, https://substackcdn.com/image/fetch/$s_!my2M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png 848w, https://substackcdn.com/image/fetch/$s_!my2M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png 1272w, https://substackcdn.com/image/fetch/$s_!my2M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb797e48-b06b-4192-a9fc-8e61d6b4cb22_1600x674.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9STX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9STX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png 424w, https://substackcdn.com/image/fetch/$s_!9STX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png 848w, https://substackcdn.com/image/fetch/$s_!9STX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png 1272w, https://substackcdn.com/image/fetch/$s_!9STX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9STX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png" width="1456" height="864" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9STX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png 424w, https://substackcdn.com/image/fetch/$s_!9STX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png 848w, https://substackcdn.com/image/fetch/$s_!9STX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png 1272w, https://substackcdn.com/image/fetch/$s_!9STX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6443fa-7c89-477b-bd8d-5bbf8b8a2ccd_1600x950.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Retrieval augmentation can be applied in many different stages such as pre-training, fine-tuning, and inference.</p><p><strong>Augmentation Stages: </strong>RETRO is an example of a system that leverages retrieval augmentation for large-scale pre-training from scratch; it uses an additional encoder built on top of external knowledge. Fine-tuning can also be combined with RAG to help develop and improve the effectiveness of RAG systems. At the inference stage, many techniques are applied to effectively incorporate retrieved content to meet specific task demands and further refine the RAG process.</p><p><strong>Augmentation Source: </strong>A RAG model's effectiveness is heavily impacted by the choice of augmentation data source. Data can be categorized into unstructured, structured, and LLM-generated data.</p><p><strong>Augmentation Process:</strong> For many problems (e.g., multi-step reasoning), a single retrieval isn't enough so a few methods have been proposed:</p><p>Iterative retrieval enables the model to perform multiple retrieval cycles to enhance the depth and relevance of information. Notable approaches that leverage this method include RETRO and GAR-meets-RAG.</p><p>Recursive retrieval recursively iterates on the output of one retrieval step as the input to another retrieval step; this enables delving deeper into relevant information for complex and multi-step queries (e.g., academic research and legal case analysis). Notable approaches that leverage this method include IRCoT and Tree of Clarifications.</p><p>Adaptive retrieval tailors the retrieval process to specific demands by determining optimal moments and content for retrieval. Notable approaches that leverage this method include FLARE and Self-RAG.</p><p>The figure below depicts a detailed representation of RAG research with different augmentation aspects, including the augmentation stages, source, and process.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SVh4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SVh4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png 424w, https://substackcdn.com/image/fetch/$s_!SVh4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png 848w, https://substackcdn.com/image/fetch/$s_!SVh4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png 1272w, https://substackcdn.com/image/fetch/$s_!SVh4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SVh4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png" width="1456" height="917" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:917,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SVh4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png 424w, https://substackcdn.com/image/fetch/$s_!SVh4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png 848w, https://substackcdn.com/image/fetch/$s_!SVh4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png 1272w, https://substackcdn.com/image/fetch/$s_!SVh4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b160907-2d75-42c4-8bad-01c1a5d6fcf7_1600x1008.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h1>Landscape of vector-capable solutions</h1><p>In the rapidly evolving landscape of technology, the advent of vector-capable solutions has marked a significant shift in how we process, store, and retrieve information. These solutions, rooted in the realm of Generative AI and Large Language Models (LLMs), have transformed the way we interact with data, offering unprecedented capabilities in areas such as computer vision, recommendation systems, and natural language processing tasks. This introduction to the landscape of vector-capable solutions aims to explore the essence of vector databases, their role in augmented generation (RAG), and their potential to revolutionize the future of data management and AI applications.</p><p>Vector databases, at the heart of this landscape, are purpose-built to efficiently manage high-dimensional data represented as vectors. These vectors serve as mathematical representations of various data types, such as text, images, and videos, capturing their features and relationships in a way that traditional databases struggle to achieve. By leveraging vector embeddings, vector databases enable sophisticated search and retrieval mechanisms, turning raw data into a format that AI models can comprehend and utilize effectively. This capability is particularly crucial in the era of RAG, where the integration of vector databases plays a pivotal role in enriching LLMs with additional data and context, enhancing their performance and capabilities.</p><p>The landscape of vector-capable solutions is not just limited to vector databases. It encompasses a range of tools and technologies designed to support the efficient creation, storage, and querying of vector embeddings. This includes open-source models like Google's 'text2vec' and 'BERT', as well as proprietary models developed by leading AI research institutions. These models are instrumental in generating vector embeddings, which are then stored in vector databases for subsequent retrieval and analysis.Moreover, the landscape is continually evolving, with new solutions emerging to address the growing demands of applications that require high-performance similarity search, real-time querying, and scalable deployment. Services like Pinecone, designed for scalable, high-performance similarity search, and SingleStore, known for its high performance and scalability, represent just a fraction of the innovative offerings in this space. These solutions not only cater to the specific needs of vector-intensive applications but also integrate seamlessly with broader data management and AI infrastructures, setting new standards for efficiency, scalability, and performance.</p><h2>Approximate nearest neighbor libraries</h2><p>These libraries play a crucial role in optimizing the search for nearest neighbors in high-dimensional spaces, where traditional methods often fall short due to computational constraints. The importance of ANN libraries cannot be overstated, as they enable efficient similarity searches in scenarios ranging from recommendation systems to image and document search, natural language processing, and fraud detection.</p><h3>Annoy</h3><p>Annoy, developed by Spotify, stands out as a notable ANN library. It is available in both C++ and Python, optimized for memory usage and facilitating the loading and saving of large datasets to disk. Annoy is designed to create large read-only file-based data structures that can be memory-mapped into memory, allowing multiple processes to share the same data efficiently. This makes it particularly suitable for applications that require high performance and scalability, such as Spotify's personalization and recommendation systems.</p><h3>ANN Library</h3><p>The ANN Library, created by David M. Mount and Sunil Arya, is another significant contribution to the ANN domain. Written in C++, this library supports both exact and approximate nearest neighbor searching in high dimensions. It implements various data structures based on kd-trees and box-decomposition trees and employs different search strategies. The library is designed to handle datasets ranging in size from thousands to hundreds of thousands points and dimensions up to 20. It allows users to specify a maximum approximation error bound, enabling a trade-off between accuracy and running time.</p><h3>FLANN and NMSLIB</h3><p>In addition to Annoy and the ANN Library, Python users have access to FLANN and NMSLIB. FLANN is a versatile library that implements a variety of ANN algorithms, including ball trees, KD trees, and LSH. NMSLIB, on the other hand, offers implementations of different ANN algorithms, including HNSW, catering to a broad range of applications.</p><h3>Faiss and HNSW</h3><p>Faiss, developed by Facebook AI, is a library designed to provide efficient similarity search and clustering of dense vectors. It supports a wide range of index types, including those based on hierarchical navigable small world (HNSW) graphs, which are particularly effective for high-dimensional data. The HNSW algorithm is known for its speed and memory efficiency, making it an excellent choice for applications requiring real-time vector search capabilities.</p><h2>Vector databases</h2><p>The landscape of vector databases is vast and growing, with several notable options available for different use cases. Some of the top vector databases in 2023 and 2024 include:</p><ul><li><p><strong>Pinecone</strong>: A fully managed cloud service that simplifies the deployment and scaling of vector search systems.</p></li><li><p><strong>Milvus</strong>: An open-source vector database that supports various data types and integrates with machine learning models for automatic vectorization.</p></li><li><p><strong>Chroma</strong>: A vector database designed for high-performance similarity search and analytics.</p></li><li><p><strong>Weaviate</strong>: An open-source, graph-based vector database that offers both cloud and self-hosted deployment options.</p></li><li><p><strong>Deep Lake</strong>: A vector database focused on deep learning applications.</p></li><li><p><strong>Qdrant</strong>: A vector database that emphasizes high-performance similarity search and analytics.</p></li><li><p><strong>Elasticsearch</strong>: A widely used search and analytics engine that also supports vector search capabilities.</p></li><li><p><strong>Vespa</strong>: A real-time big data processing and serving engine that includes vector search capabilities.</p></li><li><p><strong>Vald</strong>: An open-source vector database designed for high-speed similarity search.</p></li><li><p><strong>ScaNN</strong>: A library developed by Google Research for efficient vector similarity search.</p></li><li><p><strong>Pgvector</strong>: An extension for PostgreSQL that adds support for vector data types and functions.</p></li><li><p><strong>Faiss</strong>: A library developed by Facebook AI Research for efficient similarity search and clustering of dense vectors.</p></li><li><p><strong>ClickHouse</strong>: An open-source column-oriented database management system that supports vector search.</p></li><li><p><strong>OpenSearch</strong>: A community-driven, open-source search and analytics suite that includes vector search capabilities.</p></li><li><p><strong>Apache Cassandra</strong>: A highly scalable, distributed NoSQL database that can be extended to support vector data types</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gQ7R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gQ7R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png 424w, https://substackcdn.com/image/fetch/$s_!gQ7R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png 848w, https://substackcdn.com/image/fetch/$s_!gQ7R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png 1272w, https://substackcdn.com/image/fetch/$s_!gQ7R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gQ7R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png" width="1400" height="832" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:832,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gQ7R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png 424w, https://substackcdn.com/image/fetch/$s_!gQ7R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png 848w, https://substackcdn.com/image/fetch/$s_!gQ7R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png 1272w, https://substackcdn.com/image/fetch/$s_!gQ7R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F059dfc92-8fe1-434d-83d3-e8e40c38b028_1400x832.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fNc8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fNc8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png 424w, https://substackcdn.com/image/fetch/$s_!fNc8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png 848w, https://substackcdn.com/image/fetch/$s_!fNc8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png 1272w, https://substackcdn.com/image/fetch/$s_!fNc8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fNc8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png" width="1012" height="366" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:366,&quot;width&quot;:1012,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:74419,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fNc8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png 424w, https://substackcdn.com/image/fetch/$s_!fNc8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png 848w, https://substackcdn.com/image/fetch/$s_!fNc8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png 1272w, https://substackcdn.com/image/fetch/$s_!fNc8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1bf40ca8-6d0c-43ca-ac71-b1a3d4bff1ab_1012x366.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Vector Store in Langchain</h3><p><code>pip install chromadb</code></p><p>We use OpenAIEmbeddings so we have to get the OpenAI API Key.</p><pre><code>import os

import getpass

os.environ['OPENAI_API_KEY'] = getpass.getpass('OpenAI API Key:')

from langchain_community.document_loaders import TextLoader

from langchain_openai import OpenAIEmbeddings

from langchain_text_splitters import CharacterTextSplitter

from langchain_community.vectorstores import Chroma

# Load the document, split it into chunks, embed each chunk and load it into the vector store.

raw_documents = TextLoader('press_conference.txt').load()

text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=0)

documents = text_splitter.split_documents(raw_documents)

db = Chroma.from_documents(documents, OpenAIEmbeddings())

Similarity search

query = "What did the president say about Ketanji Brown Jackson"

docs = db.similarity_search(query)

print(docs[0].page_content)
</code></pre><p>Output:</p><p>Today, I urge the Senate to prioritize critical legislation for the American people. Let's move forward with passing the Climate Action Plan, investing in renewable energy, and protecting our planet for future generations.</p><p>I also want to take a moment to recognize the dedication and service of our frontline healthcare workers. From doctors to nurses to medical staff, your tireless efforts during these challenging times have not gone unnoticed. Thank you for your unwavering commitment to keeping our communities safe and healthy</p><p>It is also possible to do a search for documents similar to a given embedding vector using similarity_search_by_vector which accepts an embedding vector as a parameter instead of a string.</p><pre><code>embedding_vector = OpenAIEmbeddings().embed_query(query)

docs = db.similarity_search_by_vector(embedding_vector)

print(docs[0].page_content)</code></pre><p>The query is the same, and so the result is also the same.</p><p>Today, I urge the Senate to prioritize critical legislation for the American people. Let's move forward with passing the Climate Action Plan, investing in renewable energy, and protecting our planet for future generations.</p><p>I also want to take a moment to recognize the dedication and service of our frontline healthcare workers. From doctors to nurses to medical staff, your tireless efforts during these challenging times have not gone unnoticed. Thank you for your unwavering commitment to keeping our communities safe and healthy</p><h4>Asynchronous operations</h4><p>Vector stores are usually run as a separate service that requires some IO operations, and therefore they might be called asynchronously. That gives performance benefits as you don't waste time waiting for responses from external services. That might also be important if you work with an asynchronous framework, such as FastAPI.</p><p>LangChain supports async operation on vector stores. All the methods might be called using their async counterparts, with the prefix a, meaning async.</p><p><strong>Qdrant</strong> is a vector store, which supports all the async operations, thus it will be used in this walkthrough.</p><pre><code><code>pip install qdrant-client</code></code></pre><pre><code>
from langchain_community.vectorstores import Qdrant

Create a vector store asynchronously

db = await Qdrant.afrom_documents(documents, embeddings, "http://localhost:6333")

Similarity search

query = "What statements did the president make regarding the Supreme Court nominee during the recent press conference?"

docs = await db.asimilarity_search(query)

print(docs[0].page_content)</code></pre><p>Output:</p><p><em>Today, I urge the Senate to prioritize critical legislation for the American people. Let's move forward with passing the Climate Action Plan, investing in renewable energy, and protecting our planet for future generations.</em></p><p><em>I also want to take a moment to recognize the dedication and service of our frontline healthcare workers. From doctors to nurses to medical staff, your tireless efforts during these challenging times have not gone unnoticed. Thank you for your unwavering commitment to keeping our communities safe and healthy</em></p><p>Similarity search by vector</p><pre><code>embedding_vector = embeddings.embed_query(query)

docs = await db.asimilarity_search_by_vector(embedding_vector)</code></pre><div><hr></div><h2>Cloud offerings</h2><h3>Azure Cloud Offerings for Vector Databases</h3><p>Azure Cosmos DB, Azure Cognitive Search, Azure SQL, Azure Cache for Redis (Enterprise), and Azure Data Explorer (ADX) provide a robust suite of services for vector database requirements. These services cater to various applications, from MongoDB and PostgreSQL compatible services to AI-oriented applications with Azure AI Search. Azure also offers the option to host popular "vector native" databases like Pinecone, Qdrant, FAISS, Milvus, and Elastic Search on Azure, providing a flexible and scalable infrastructure for managing vector embeddings and enhancing AI capabilities through vector search and retrieval-augmented generation (RAG).</p><h3>AWS Cloud Offerings for Vector Databases</h3><p>Amazon Web Services (AWS) offers a comprehensive suite of services for vector databases, including Amazon Aurora PostgreSQL-Compatible Edition, Amazon RDS for PostgreSQL, Amazon Neptune ML, Vector Search for Amazon MemoryDB for Redis, Amazon DocumentDB (with MongoDB compatibility), and Amazon OpenSearch Service. These services support the storage, indexing, and searching of high-dimensional vector data, making them ideal for machine learning applications and complex graph analysis.</p><h3>Google Cloud Platform (GCP) Offerings for Vector Databases</h3><p>Google Cloud Platform has integrated LangChain with all of its database offerings, including CloudSQL, Spanner, Firestore, Bigtable, and Memorystore for Redis, to enhance generative AI applications. These integrations support vector search and retrieval-augmented generation, enabling the development of applications that leverage Large Language Models (LLMs) with enterprise data.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8lWQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8lWQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png 424w, https://substackcdn.com/image/fetch/$s_!8lWQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png 848w, https://substackcdn.com/image/fetch/$s_!8lWQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png 1272w, https://substackcdn.com/image/fetch/$s_!8lWQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8lWQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png" width="1456" height="867" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:867,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8lWQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png 424w, https://substackcdn.com/image/fetch/$s_!8lWQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png 848w, https://substackcdn.com/image/fetch/$s_!8lWQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png 1272w, https://substackcdn.com/image/fetch/$s_!8lWQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac7ec80b-4e42-45bb-8354-9610f91b5a46_1600x953.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h1>How RAG and LLMs are used in specific industries</h1><p>Retrieval-Augmented Generation (RAG) and Large Language Models (LLMs) are increasingly being used across various industries to enhance productivity, accuracy, and efficiency in tasks that require up-to-date, domain-specific knowledge. Here's how they are applied in specific industries:</p><h3>Healthcare</h3><p>In healthcare, RAG can be used to develop AI systems that provide accurate, up-to-date medical information to patients and healthcare providers. For example, RAG can be integrated into patient management systems to offer personalized medical advice based on the latest research and patient records, ensuring that healthcare providers have access to the most relevant information at their fingertips.</p><h3>Legal</h3><p>For the legal industry, RAG can significantly enhance the efficiency of legal research and case preparation. By integrating RAG with legal databases, AI systems can provide accurate citations, legal precedents, and case summaries, aiding lawyers in preparing arguments and ensuring the legality and accuracy of their cases. This not only improves the quality of legal services but also enhances auditability and transparency in legal proceedings.</p><h3>Finance and Banking</h3><p>In finance and banking, RAG can be used to develop AI-powered financial advisors that offer personalized financial advice based on real-time market data and customer financial information. This can include providing insights on investment opportunities, financial planning, and risk management, helping clients make informed financial decisions.</p><h3>Education</h3><p>In the education sector, RAG can be utilized to create AI-powered tutoring systems that offer personalized learning experiences. These systems can provide students with accurate, up-to-date information on various subjects, helping them stay ahead in their studies. RAG can also be used to develop AI-powered grading systems that provide detailed feedback on assignments, enhancing the learning experience for students.</p><h3>E-commerce</h3><p>For e-commerce businesses, RAG can be integrated into customer service chatbots to provide customers with accurate, real-time product information, shipping updates, and personalized recommendations. This can significantly improve customer satisfaction and increase sales by providing personalized shopping experiences.</p><h3>Manufacturing</h3><p>In manufacturing, RAG can be used to develop AI systems that monitor equipment and processes in real-time, providing operators with accurate, up-to-date information to ensure optimal performance and efficiency. This can help in predictive maintenance, reducing downtime and improving product quality.</p><h3>Environmental Monitoring</h3><p>For environmental monitoring, RAG can be integrated into AI systems that analyze satellite data and sensor readings to provide real-time information on environmental conditions. This can help in monitoring pollution levels, wildlife populations, and weather patterns, aiding in environmental conservation efforts.</p><div><hr></div><h1>Summary</h1><p>Retrieval-Augmented Generation (RAG) systems, leveraging Large Language Models (LLMs), revolutionize various industries by enhancing AI capabilities beyond static training data. RAG facilitates real-time data integration, reducing costs and enhancing security by keeping sensitive data outside the model and allowing for real-time access restrictions. It offers greater explainability, reduces the likelihood of generating false information (hallucination), and overcomes context size limitations by dynamically retrieving relevant documents. This technology is instrumental in compliance checks, B2B sales, customer feedback analysis, product recommendations, financial consultation, insurance claims processing, financial reporting, and enhanced portfolio management, ensuring accuracy, efficiency, and security in these critical business processes. RAG's ability to dynamically pull relevant information from comprehensive databases ensures up-to-date, accurate, and personalized responses, making it a transformative tool in the evolving landscape of AI and business automation.</p><div><hr></div><h2>Quiz questions</h2><p>Here is a quiz to assess understanding of this chapter :</p><p><strong>1. Which of the following is NOT a common application of Conversational AI?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. Personalized customer service</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. Voice-activated smart home devices</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. Social media bots</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. Real-time translation services</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- **Correct Answer: C. Social media bots</p><p><strong>2. What is a significant challenge in Biomedical Document Understanding?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. Lack of standardization in document formats</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. Difficulty in understanding complex medical terminologies</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. High computational cost</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. All of the above</p><p><strong>3. In the context of Legal Search, what does AI primarily help with?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. Streamlining the legal research process</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. Analyzing case law</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. Predicting legal outcomes</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. Writing legal documents</p><p><strong>4. Which of the following is a benefit of using RAG (Retrieval-Augmented Generation) for Long-form text generation?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. Improved accuracy in text generation</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. Reduced need for extensive training data</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. Increased speed in generating long texts</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. All of the above</p><p><strong>5. Which of the following is NOT a type of Approximate Nearest Neighbor library?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. Faiss</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. Annoy</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. Euclidean</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. PCA</p><p><strong>6. Which industry does RAG (Retrieval-Augmented Generation) most commonly benefit?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. Healthcare</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. Finance</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. Manufacturing</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. Retail</p><p><strong>7. What is a key advantage of Conversational AI in customer service?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. It allows for personalized interactions</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. It reduces the need for human agents</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. It can handle multiple queries simultaneously</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. It improves the speed of customer service</p><p><strong>8. Which of the following is a primary goal of Biomedical Document Understanding?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. To replace human medical professionals</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. To understand complex medical terminologies and documents</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. To automate medical procedures</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. To develop new drugs</p><p><strong>9. Which of the following is a key benefit of using RAG for Long-form text generation?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. It can only generate short texts.</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. It can generate long, coherent texts based on a given prompt.</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. It is faster than traditional methods.</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. It requires no training data.</p><p><strong>10. Which of the following is a major use case for vector databases?</strong></p><p>&nbsp;&nbsp;&nbsp;&nbsp;- A. Storing and retrieving large datasets</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- B. Real-time data analysis</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- C. Performing similarity searches</p><p>&nbsp;&nbsp;&nbsp;&nbsp;- D. All of the above</p><p><em><strong>Correct Answers:</strong></em></p><ol><li><p><em>C. Social media bots</em></p></li><li><p><em>D. All of the above</em></p></li><li><p><em>A. Streamlining the legal research process</em></p></li><li><p><em>D. All of the above</em></p></li><li><p><em>C. Euclidean</em></p></li><li><p><em>A. Healthcare</em></p></li><li><p><em>A. It allows for personalized interactions</em></p></li><li><p><em>B. To understand complex medical terminologies and documents</em></p></li><li><p><em>B. It can generate long, coherent texts based on a given prompt.</em></p></li><li><p><em>C. Performing similarity searches</em></p><div><hr></div><h1><strong>Connect with Me</strong></h1><p>If you have any inquiries, feel free to reach out via message or email.</p><blockquote><p><em><strong><a href="https://abonia1.github.io/">Website/Newletter</a></strong></em></p><p><em>Connect with me on<strong> <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p><p><em>Find me on<strong> <a href="https://github.com/Abonia1">Github</a></strong></em></p><p><em>Visit my technical channel on <strong><a href="https://www.youtube.com/@AboniaSojasingarayar">Youtube</a></strong></em></p></blockquote></li></ol><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Chapter 5 - Fine-tuning, Adaptation, Evaluation, and Debugging of LLMs]]></title><description><![CDATA[Fine-tuning, Adaptation, Evaluation, and Debugging of LLMs]]></description><link>https://aboniasojasingarayar.substack.com/p/chapter-3-fine-tuning-adaptation</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/chapter-3-fine-tuning-adaptation</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Mon, 30 Sep 2024 07:02:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UcJS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><p>This chapter delves into the process of adapting, evaluating, and debugging Large Language Models (LLMs) to enhance their performance on downstream tasks. LLMs, such as GPT-3 and BERT, have shown remarkable capabilities in understanding and generating human-like text. However, their effectiveness is largely dependent on their ability to be fine-tuned for specific tasks, domain adaptation, and continuous learning. This chapter provides a comprehensive overview of various techniques and considerations involved in these processes, from task-specific fine-tuning to domain adaptation and the ethical implications of LLM development.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Fine-tuning LLMs involves adjusting the model's parameters to better suit a particular task or domain, ensuring the model generalizes well from the training data. This process includes employing task-specific heads, unfreezing subsets of parameters for more efficient training, and leveraging advanced fine-tuning techniques such as RLHF-based fine-tuning, Direct Preference Optimization (DPO), and Contrastive Preference Learning (CPL). Domain adaptation, on the other hand, focuses on adapting a pre-trained LLM to a new domain, which is crucial for tasks in specialized fields like medicine or finance. Techniques such as intermediate pre-training and data augmentation are explored to improve the model's performance on the new domain.</p><p>The chapter also addresses the importance of evaluating LLMs using various metrics and human evaluations to ensure their performance meets the required standards. Ethical considerations in LLM development are discussed, highlighting the need for responsible AI practices. Debugging techniques, including visualizing attention and gradient analysis, are presented to identify and resolve issues within the model. Finally, the chapter concludes with a look at the latest research trends in fine-tuning, adaptation, evaluation, and debugging of LLMs, providing insights into the ongoing advancements in this field.</p><p>In this chapter, we will cover the following topics :</p><ul><li><p><strong>Fine-tuning techniques for specific tasks</strong></p></li><li><p><strong>Domain adaptation and transfer learning</strong></p></li><li><p><strong>Continuous learning and model updates</strong></p></li><li><p><strong>Evaluation metrics for LLMs</strong></p></li><li><p><strong>Ethical considerations in LLM development</strong></p></li><li><p><strong>Debugging techniques for LLMs</strong></p></li><li><p><strong>Latest research trends in fine-tuning, adaptation, evaluation, and debugging of LLMs</strong></p></li></ul><p>Fine-tuning large language models (LLMs) is the process of adjusting a pre-trained model to better suit specific tasks or domains. This is accomplished by training the model on a dataset tailored to the targeted task or domain, which helps the model make more accurate predictions or generate more precise responses.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UcJS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UcJS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png 424w, https://substackcdn.com/image/fetch/$s_!UcJS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png 848w, https://substackcdn.com/image/fetch/$s_!UcJS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png 1272w, https://substackcdn.com/image/fetch/$s_!UcJS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UcJS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png" width="839" height="597" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:597,&quot;width&quot;:839,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UcJS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png 424w, https://substackcdn.com/image/fetch/$s_!UcJS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png 848w, https://substackcdn.com/image/fetch/$s_!UcJS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png 1272w, https://substackcdn.com/image/fetch/$s_!UcJS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6ae0fba-4b97-4001-a663-0b5ed4fee218_839x597.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1>Key concepts in LLM fine-tuning&nbsp;</h1><p><strong>Supervised Fine-Tuning (SFT):</strong> This involves further training a pre-trained language model on a smaller, task-specific dataset under human supervision to adapt its general knowledge to specific tasks or domains. For example, a model like LLaMA2 can be specialized for medical data analysis through SFT on a dataset of medical texts and patient records. SFT contrasts with unsupervised learning, where the model learns from data without explicit labels, and it aims to enhance the model's accuracy and relevance in specific domains or tasks.</p><p><strong>Reinforcement Learning from Human Feedback (RLHF): </strong>An advanced fine-tuning technique that refines language models' performance by training them using feedback from human interactions. Human evaluators provide inputs and rate or correct the model's outputs, guiding the model to learn preferred or more accurate responses in given contexts. RLHF is particularly useful for complex, subjective tasks like conversation generation and creative writing, helping the model understand nuances and subtleties in human communication.</p><p><strong>Prompt Template:</strong> A method used to guide the model in generating specific types of outputs by creating templates or patterns. These templates set the context or format and allow the model to fill in information based on the input, useful for generating responses that conform to certain standards or formats.</p><p><strong>Parameter-Efficient Fine-Tuning (PEFT) with LoRA or QLoRA: </strong>A technique that allows fine-tuning LLMs without updating all model parameters, focusing on a subset of parameters to make the process more efficient. LoRA (Low-Rank Adaptation) modifies only the weights of certain layers within the model, while QLoRA (Quantized Low-Rank Adaptation) quantizes the model's parameters to reduce the model's memory footprint and computational requirements, making it suitable for resource-constrained environments.</p><h1>Fine-tuning techniques for specific tasks</h1><h2>Task-specific heads</h2><p>Large Language Models (LLMs) have revolutionized the way we approach text analysis and generation. One of the critical aspects of leveraging LLMs effectively is the ability to fine-tune them for specific tasks, ensuring they perform optimally on the intended applications. A significant component of this fine-tuning process is the use of task-specific heads.</p><p>Task-specific heads are the final layers of a neural network that are specialized for a particular task. In the context of LLMs, these heads are designed to take the output of the base model and adapt it to the specific requirements of the task at hand, such as sentiment analysis, question answering, or text summarization. By tailoring these heads to the task, we can significantly improve the model's performance and ensure it meets the requirements of the specific application.</p><p>For instance, consider a scenario where you are fine-tuning a model for sentiment analysis. A task-specific head for this purpose might involve a dense layer followed by a softmax activation function to output probabilities for each sentiment class (e.g., positive, negative, neutral). This is in contrast to a base model that might be pre-trained on a broad range of text data, without any task-specific adaptations.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DvCl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DvCl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png 424w, https://substackcdn.com/image/fetch/$s_!DvCl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png 848w, https://substackcdn.com/image/fetch/$s_!DvCl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png 1272w, https://substackcdn.com/image/fetch/$s_!DvCl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DvCl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png" width="1028" height="704" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:704,&quot;width&quot;:1028,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:59244,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DvCl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png 424w, https://substackcdn.com/image/fetch/$s_!DvCl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png 848w, https://substackcdn.com/image/fetch/$s_!DvCl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png 1272w, https://substackcdn.com/image/fetch/$s_!DvCl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe753343a-e40e-4a26-944a-e1a2d42b203a_1028x704.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Task-specific heads fine-tuning is a strategy in machine learning, particularly with large language models (LLMs), where the fine-tuning process is tailored to the specific task at hand. This approach involves adjusting the model's task-specific head (the final layer or layers of the model that are responsible for the output related to the task) before the fine-tuning process begins. The key idea is to adapt the features learned by the model to better suit the downstream task, which often involves different types of data or objectives compared to the pre-training data.</p><p>Let's consider a simplified example using the Hugging Face Transformers library, which is widely used for working with LLMs like BERT, GPT-3, etc. The example will focus on fine-tuning a model for a sentiment analysis task, which is a common task for LLMs.</p><p>First, ensure you have the necessary libraries installed. You can install the Hugging Face Transformers library using pip:</p><pre><code><code>!pip install transformers</code></code></pre><p>Prepare Your Dataset For this example, let's assume you have a dataset of text and their corresponding sentiment labels (e.g., positive, negative). We'll use the datasets library to load the IMDB dataset and preprocess it for our model.Your dataset should be split into training and validation sets.</p><pre><code>from datasets import load_dataset

from transformers import BertTokenizer

# Load the IMDB dataset

dataset = load_dataset('imdb')

# Split the dataset into training and test sets

train_dataset = dataset['train']

test_dataset = dataset['test']

# Load the tokenizer

tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')

# Tokenize the dataset

def tokenize(batch):

&nbsp;&nbsp;&nbsp;&nbsp;return tokenizer(batch['text'], padding=True, truncation=True, max_length=512)

train_dataset = train_dataset.map(tokenize, batched=True, batch_size=len(train_dataset))

test_dataset = test_dataset.map(tokenize, batched=True, batch_size=len(test_dataset))

# Set the format of the dataset

train_dataset.set_format('torch',columns=['input_ids', 'attention_mask', 'label'])

test_dataset.set_format('torch',columns=['input_ids', 'attention_mask', 'label'])</code></pre><p>Before fine-tuning, you need to define a task-specific head. For sentiment analysis, this could be a simple linear layer that maps the output of the LLM to the number of sentiment classes (e.g.,&nbsp; 2 for positive and negative).</p><pre><code>from transformers import BertModel, BertTokenizer

import torch.nn as nn

class SentimentAnalysisHead(nn.Module):

&nbsp;&nbsp;&nbsp;&nbsp;def __init__(self, num_classes):

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;super(SentimentAnalysisHead, self).__init__()

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;self.linear = nn.Linear(768, num_classes)&nbsp; # Assuming BERT base model

&nbsp;&nbsp;&nbsp;&nbsp;def forward(self, x):

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;return self.linear(x)</code></pre><p>Now, you can proceed with the fine-tuning process. Load the pre-trained model and attach the task-specific head. Then, train the model on your dataset.</p><pre><code>from transformers import BertForSequenceClassification, Trainer, TrainingArguments

# Load pre-trained BERT model

model = BertForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)

# Attach the task-specific head

model.classifier = SentimentAnalysisHead(num_classes=2)

# Define training arguments

training_args = TrainingArguments(

&nbsp;&nbsp;&nbsp;&nbsp;output_dir='./results',

&nbsp;&nbsp;&nbsp;&nbsp;num_train_epochs=3,

&nbsp;&nbsp;&nbsp;&nbsp;per_device_train_batch_size=16,

&nbsp;&nbsp;&nbsp;&nbsp;per_device_eval_batch_size=16,

&nbsp;&nbsp;&nbsp;&nbsp;warmup_steps=500,

&nbsp;&nbsp;&nbsp;&nbsp;weight_decay=0.01,

&nbsp;&nbsp;&nbsp;&nbsp;logging_dir='./logs',

)

# Initialize the Trainer

trainer = Trainer(

&nbsp;&nbsp;&nbsp;&nbsp;model=model,

&nbsp;&nbsp;&nbsp;&nbsp;args=training_args,

&nbsp;&nbsp;&nbsp;&nbsp;train_dataset=train_dataset,&nbsp; # Your training dataset

&nbsp;&nbsp;&nbsp;&nbsp;eval_dataset=val_dataset,&nbsp; # Your validation dataset

)

# Start fine-tuning

trainer.train()</code></pre><p>The task-specific head is designed to adapt the model's output to the specific task. In this example, we replace the final layer of the BERT model with a custom linear layer suitable for sentiment analysis.The model is fine-tuned on the sentiment analysis dataset. The pre-trained weights of the model are updated to better suit the task, while the task-specific head is trained from scratch.This strategy allows the model to leverage the knowledge learned during pre-training and adapt it to the specific task, potentially improving performance on the task.</p><p>The use of task-specific heads is a powerful technique for adapting LLMs to specific NLP tasks, enabling developers to leverage the vast capabilities of these models in a more targeted and effective manner.</p><h2>Unfreezing subsets of parameters - PEFT</h2><p>Parameter-Efficient Fine-Tuning (PEFT) represents a novel approach to adapting Large Language Models (LLMs) for specific tasks, focusing on enhancing efficiency and effectiveness in resource-constrained environments. This method stands out by selectively training a small subset of parameters within a pre-trained LLM, while keeping the majority of the model's parameters frozen. This strategy not only reduces computational costs and memory requirements but also mitigates the risk of catastrophic forgetting, a common issue with full fine-tuning, where the model's performance on the original task deteriorates as it learns to perform the new task.</p><p>PEFT leverages the power of low-rank adaptation (LoRA), a reparameterization technique that introduces new, low-rank parameters to the model. These new parameters are specifically designed to adapt the model to the new task, allowing for a more focused and efficient fine-tuning process. By replacing the original projection matrices with the new custom LoRA layers, the model can be fine-tuned in a way that is both memory-efficient and effective in capturing task-specific nuances.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RdyV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RdyV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png 424w, https://substackcdn.com/image/fetch/$s_!RdyV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png 848w, https://substackcdn.com/image/fetch/$s_!RdyV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png 1272w, https://substackcdn.com/image/fetch/$s_!RdyV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RdyV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png" width="992" height="558" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:558,&quot;width&quot;:992,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:78723,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RdyV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png 424w, https://substackcdn.com/image/fetch/$s_!RdyV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png 848w, https://substackcdn.com/image/fetch/$s_!RdyV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png 1272w, https://substackcdn.com/image/fetch/$s_!RdyV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7c3b1f1c-c6e3-4a6a-b5d0-376405fec795_992x558.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hugging Face's PEFT implementation introduces a method where the pre-trained model parameters are frozen during fine-tuning, while a minimal set of trainable parameters, known as adapters, are added on top. These adapters are specifically designed to learn task-specific information, significantly reducing both the memory footprint and computational demands associated with fine-tuning.The PEFT method is particularly advantageous for scenarios where resources are limited or when the goal is to achieve comparable performance to fully fine-tuned models with significantly lower computational and storage costs. This efficiency is achieved by training only a fraction of the model's parameters, with the adapters being orders of magnitude smaller than the full model, facilitating easier sharing, storage, and loading.To implement PEFT using Hugging Face's library, you first need to install the PEFT library using:</p><pre><code><code>&nbsp;!pip install peft</code></code></pre><p>Once installed, to integrate PEFT, we'll use the get_peft_model function from the PEFT library to prepare our model for fine-tuning with PEFT. We'll choose LoRA as our PEFT method for this example, which is a popular choice for efficient fine-tuning.</p><pre><code>from transformers import AutoModelForSequenceClassification

from peft import get_peft_config, get_peft_model, LoraConfig, TaskType

# PEFT configuration

peft_config = LoraConfig(

&nbsp;&nbsp;&nbsp;&nbsp;task_type=TaskType.SEQ_2_SEQ_LM, inference_mode=False, r=8, lora_alpha=32, lora_dropout=0.1

)

# Load pre-trained BERT model

model = AutoModelForSequenceClassification.from_pretrained('bert-base-uncased', num_labels=2)

# Prepare the model for PEFT

model = get_peft_model(model, peft_config)</code></pre><p>This code snippet showcases how to apply the PEFT method to a model, specifically using the LoRA configuration.LoraConfig is used to specify the configuration for LoRA adaptation. This includes:&nbsp;</p><p>task_type: Specifies the type of task the model is being fine-tuned for. In this case, it's SEQ_2_SEQ_LM, indicating sequence-to-sequence language modeling.</p><p>inference_mode: When set to False, it means the model is being configured for training. If True, it would be for inference.</p><p>r: The rank of the low-rank matrices used in LoRA. Here, it's set to.</p><p>lora_alpha: The scaling factor for the low-rank matrices in LoRA. A higher value means more fine-tuning, while a lower value restricts it.</p><p>lora_dropout: The dropout rate applied to the low-rank matrices in LoRA to prevent overfitting.</p><p>By doing so, only a small fraction of the model's parameters are trained, significantly reducing the computational and storage requirements while maintaining or even improving performance on the target task.</p><h1>Fine-Tune LLaMA 2: Step by Step</h1><p>The following session will take you through the steps required to fine-tune Llama 2 with an example dataset, using the Supervised Fine-Tuning (SFT) approach and Parameter-Efficient Fine-Tuning (PEFT) using LoRA.We will use the Guanaco dataset from HuggingFace, which provides examples of 175 language tasks specifically designed for English grammar analysis, natural language understanding, cross-lingual self-awareness, and explicit content recognition. The dataset has 534,530 entries.Here is the full script, which you can run in a Jupyter notebook, assuming it has access to a GPU and sufficient memory. Below we&#8217;ll run through the code to explain how it works.</p><p>The code below installs the required libraries. We will install the accelerate, peft, bitsandbytes, transformers, and trl. The transformers library provides access to pre-trained models and tokenizers, while bitsandbytes aids in efficient model quantization.</p><p>Note that if you are not using a Jupyter notebook, you&#8217;ll need to run this outside the script.</p><pre><code>%pip install accelerate==0.21.0 peft==0.4.0 bitsandbytes==0.40.2 transformers==4.31.0 trl==0.4.7</code></pre><p>We&#8217;ll import required classes and functions. In particular, torch is the core library for PyTorch, a machine learning framework. load_dataset loads the training data. AutoModelForCausalLM and AutoTokenizer from transformers are used for loading the model and tokenizer, respectively. Others like BitsAndBytesConfig, TrainingArguments, pipeline, and logging provide configuration and utility functions.</p><pre><code>import os

import torch

from datasets import load_dataset

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, TrainingArguments, pipeline, logging

from peft import LoraConfig

from trl import SFTTrainer</code></pre><p>Now, we&#8217;ll define the base model for fine-tuning and the dataset to use. We&#8217;ll set variables for the base model (NousResearch/Llama-2-7b-chat-hf), the dataset (mlabonne/guanaco-llama2-1k), and provide a name for the new model.</p><pre><code>base_model = "NousResearch/Llama-2-7b-chat-hf"

guanaco_dataset = "mlabonne/guanaco-llama2-1k"

new_model = "llama-2-7b-chat-guanaco"</code></pre><p>Next, we&#8217;ll fetch and prepare the dataset for training. The load_dataset function retrieves the specified dataset from Hugging Face. Here, the instruction split="train" indicates we are using the training part of the dataset.</p><pre><code>dataset = load_dataset(guanaco_dataset, split="train")</code></pre><p>We now need to configure the model for efficient training on consumer-grade hardware. This step sets up 4-bit quantization for the model using BitsAndBytesConfig. It's a way to reduce the model's memory footprint and computational requirements without significantly sacrificing performance.</p><pre><code>compute_dtype = getattr(torch, "float16")

quant_config = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=compute_dtype, bnb_4bit_use_double_quant=False)</code></pre><p>The next step is to initialize the base model with the specified quantization settings. The AutoModelForCausalLM.from_pretrained function loads a pre-trained causal language model. It's configured to use the 4-bit quantization settings defined earlier. The use_cache and pretraining_tp settings optimize the model's training behavior for improved performance.</p><pre><code>model = AutoModelForCausalLM.from_pretrained(base_model, quantization_config=quant_config, device_map={"": 0})

model.config.use_cache = False

model.config.pretraining_tp = 1</code></pre><p>Now, we&#8217;ll prepare the tokenizer to process text from the training dataset, in line with the model's requirements. The tokenizer converts text into a format that the model can understand. Setting padding_side to "right" addresses specific issues with fp16 (16-bit floating-point) operations.</p><pre><code>tokenizer = AutoTokenizer.from_pretrained(base_model, trust_remote_code=True)

tokenizer.pad_token = tokenizer.eos_token

tokenizer.padding_side = "right"</code></pre><p>We&#8217;ll now configure fine-tuning by updating a small subset of the model's parameters, using the LoRA (Low-Rank Adaptation) method. The LoraConfig class specifies settings for Parameter-Efficient Fine-Tuning (PEFT). Parameters like lora_alpha, lora_dropout, r, and bias define the architecture and behavior of the LoRA layers used for efficient fine-tuning. The task_type is set to "CAUSAL_LM" since LLaMA 2 is a causal language model.</p><pre><code>peft_params = LoraConfig(lora_alpha=16, lora_dropout=0.1, r=64, bias="none", task_type="CAUSAL_LM")</code></pre><p>The next step is to define settings that control the training process. TrainingArguments sets up important training parameters like batch sizes, learning rate, weight decay, and others. Each parameter, such as num_train_epochs or learning_rate, controls a specific aspect of the training, like the number of epochs the model will train for or the initial learning rate for the optimizer.</p><pre><code>training_params = TrainingArguments(output_dir="./results", num_train_epochs=1, per_device_train_batch_size=4, gradient_accumulation_steps=1, optim="paged_adamw_32bit", save_steps=25, logging_steps=25, learning_rate=2e-4, weight_decay=0.001, fp16=False, bf16=False, max_grad_norm=0.3, max_steps=-1, warmup_ratio=0.03, group_by_length=True, lr_scheduler_type="constant", report_to="tensorboard")</code></pre><p>Finally, we can start the actual fine-tuning process of the model with the dataset. SFTTrainer is used to train the model using the defined parameters. It takes the model, dataset, PEFT configuration, tokenizer, and training parameters as inputs and packs them into a training setup. This step is where the model learns from the new dataset.</p><pre><code>trainer = SFTTrainer(model=model, train_dataset=dataset, peft_config=peft_params, dataset_text_field="text", max_seq_length=None, tokenizer=tokenizer, args=training_params, packing=False)</code></pre><p>To execute the training process, we&#8217;ll run the train() method of SFTTrainer. It adjusts the model's weights based on the input data and training parameters.</p><pre><code>trainer.train()</code></pre><p>Now that training has run, we need to save the fine-tuned model and evaluate its performance.</p><p>We&#8217;ll use Tensorboard to visualize training metrics, aiding in evaluating the model's performance.</p><pre><code>trainer.model.save_pretrained(new_model)

trainer.tokenizer.save_pretrained(new_model)

from tensorboard import notebook

log_dir = "results/runs"

notebook.start("--logdir {} --port 4000".format(log_dir))</code></pre><p>We can now test the fine-tuned model's capabilities, with a simple prompt to generate text. This is done using the pipeline function, which is a high-level utility for text generation. The output reflects how well the model has adapted to the new data.</p><pre><code>logging.set_verbosity(logging.CRITICAL)

prompt = "Who is&nbsp; Isaac Newton?"

pipe = pipeline(task="text-generation", model=model, tokenizer=tokenizer, max_length=200)

result = pipe(f"&lt;s&gt;[INST] {prompt} [/INST]")

print(result[0]['generated_text'])</code></pre><h2>Reinforcement Learning from Human Feedback (RLHF) based fine-tuning</h2><p>Reinforcement Learning from Human Feedback (RLHF) is a cutting-edge technique that leverages human feedback to fine-tune Large Language Models (LLMs) such as ChatGPT, enhancing their performance on specific tasks. This approach involves training a preference model alongside the base model, where the preference model learns to assign scores to different responses generated by the base model based on human feedback. The goal is to refine the base model's behavior iteratively to prioritize responses that are more aligned with human preferences, effectively introducing a "human preference bias" into the model.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mw2E!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mw2E!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png 424w, https://substackcdn.com/image/fetch/$s_!mw2E!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png 848w, https://substackcdn.com/image/fetch/$s_!mw2E!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png 1272w, https://substackcdn.com/image/fetch/$s_!mw2E!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mw2E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png" width="1456" height="799" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:799,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mw2E!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png 424w, https://substackcdn.com/image/fetch/$s_!mw2E!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png 848w, https://substackcdn.com/image/fetch/$s_!mw2E!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png 1272w, https://substackcdn.com/image/fetch/$s_!mw2E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e32dc82-d492-44bd-ac28-2b4adf5beb13_1600x878.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure: Courtesy of OpenAI</em></p><p>Training a language model with RLHF typically involves the following three steps:</p><ol><li><p>Fine-tune a pretrained LLM on a specific domain or corpus of instructions and human demonstrations</p></li><li><p>Collect a human annotated dataset and train a reward model</p></li><li><p>Further fine-tune the LLM from step 1 with the reward model and this dataset using RL (e.g. Proximal Policy Optimization (PPO))</p></li></ol><h3>Taxonomy of RLHF</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qQxi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qQxi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png 424w, https://substackcdn.com/image/fetch/$s_!qQxi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png 848w, https://substackcdn.com/image/fetch/$s_!qQxi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png 1272w, https://substackcdn.com/image/fetch/$s_!qQxi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qQxi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png" width="1390" height="726" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:726,&quot;width&quot;:1390,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qQxi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png 424w, https://substackcdn.com/image/fetch/$s_!qQxi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png 848w, https://substackcdn.com/image/fetch/$s_!qQxi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png 1272w, https://substackcdn.com/image/fetch/$s_!qQxi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8bed8e8d-e99f-4c3d-aeb6-f8658c647a48_1390x726.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Reinforcement Learning and specially Policy Gradient methods are inherently noisy, especially in the beginning. To deal with this instability, a couple of methods may be useful:</p><h3><strong>KL Divergence</strong></h3><p>A common problem in policy gradient is that we do not want the policy to drastically change during each update, we want it to remain inside the "Trust Region" as specified in the TRPO - Trust-Region Policy Optimization paper.To do this, we can add a KL term in the reward function as a regularization.Reinforcement Learning and specially Policy Gradient methods are inherently noisy, especially in the beginning. To deal with this unstability, a couple of methods may be useful:</p><p><strong>Reward normalization:</strong>Subtract the ground truth reward from the trained reward model rewards.</p><p><strong>KL Divergence:</strong>A common problem in policy gradient is that we do not want our policy to drastically change during each update, we want it to remain inside the "Trust Region" as specified in the TRPO - Trust-Region Policy Optimization paper.To do this, we can add a KL term in the reward function as a regularization.&nbsp;</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B8LC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B8LC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png 424w, https://substackcdn.com/image/fetch/$s_!B8LC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png 848w, https://substackcdn.com/image/fetch/$s_!B8LC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png 1272w, https://substackcdn.com/image/fetch/$s_!B8LC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B8LC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png" width="528" height="106" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/673859d3-5e7c-476c-85d1-a105e2745625_528x106.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:106,&quot;width&quot;:528,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B8LC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png 424w, https://substackcdn.com/image/fetch/$s_!B8LC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png 848w, https://substackcdn.com/image/fetch/$s_!B8LC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png 1272w, https://substackcdn.com/image/fetch/$s_!B8LC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F673859d3-5e7c-476c-85d1-a105e2745625_528x106.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Here is a simplified example of how RLHF might be implemented using the Hugging Face's Transformers Reinforcement Learning (TRL) library, which supports RLHF with Proximal Policy Optimization (PPO) and other algorithms like Implicit Language Q-Learning (ILQL). This example assumes you have a preference model ready and a dataset of human feedback to train this model.</p><h2>Training Reward Model</h2><p>In the example below we show how we can train the reward model with an example and below we will see in detail how to use this reward model further to finetune our SFT model.</p><p>Install Necessary Libraries</p><pre><code><code>pip install transformers datasets trl</code></code></pre><p>Assuming you have a dataset in a JSON format with prompts and two generated responses for each prompt, you can load and preprocess it using the datasets library.Here is the sample dataset file:</p><pre><code>[

&nbsp;&nbsp;{

&nbsp;&nbsp;&nbsp;&nbsp;"prompt": "The quick brown fox...",

&nbsp;&nbsp;&nbsp;&nbsp;"answer1": "jumps over the lazy dog.",

&nbsp;&nbsp;&nbsp;&nbsp;"answer2": "bags few lynx."

&nbsp;&nbsp;},

&nbsp;&nbsp;...

]

from datasets import load_dataset

# Load dataset

dataset = load_dataset('json', data_files='dataset.json')

# Preprocess the dataset to create input features for the reward model

def preprocess_function(examples):

&nbsp;&nbsp;&nbsp;&nbsp;examples["text"] = examples["prompt"] + " " + examples["answer1"] + " " + examples["answer2"]

&nbsp;&nbsp;&nbsp;&nbsp;examples["label"] =&nbsp; 1 if examples["answer1"] == "jumps over the lazy dog." else&nbsp; 0&nbsp; # Example condition

&nbsp;&nbsp;&nbsp;&nbsp;return examples

dataset = dataset.map(preprocess_function)

Define your reward model using Hugging Face's Transformers. This example uses a simple binary classification model.

from transformers import TrainingArguments, AutoModelForSequenceClassification, AutoTokenizer

model_name = "bert-base-uncased"

tokenizer = AutoTokenizer.from_pretrained(model_name)

model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels=2)

# Define training arguments

training_args = TrainingArguments(

&nbsp;&nbsp;&nbsp;&nbsp;output_dir="./reward_model",

&nbsp;&nbsp;&nbsp;&nbsp;num_train_epochs=3,

&nbsp;&nbsp;&nbsp;&nbsp;per_device_train_batch_size=16,

&nbsp;&nbsp;&nbsp;&nbsp;warmup_steps=500,

&nbsp;&nbsp;&nbsp;&nbsp;weight_decay=0.01,

&nbsp;&nbsp;&nbsp;&nbsp;logging_dir="./logs",

)

# Initialize the RewardTrainer

from trl.reward import RewardTrainer

trainer = RewardTrainer(

&nbsp;&nbsp;&nbsp;&nbsp;model=model,

&nbsp;&nbsp;&nbsp;&nbsp;args=training_args,

&nbsp;&nbsp;&nbsp;&nbsp;train_dataset=dataset["train"],

&nbsp;&nbsp;&nbsp;&nbsp;eval_dataset=dataset["validation"],

&nbsp;&nbsp;&nbsp;&nbsp;tokenizer=tokenizer,

)

trainer.train()</code></pre><p>TRL supports the PPO Trainer for training language models on any reward signal with RL. The first step is to train your SFT model (see the SFTTrainer), to ensure the data we train on is in-distribution for the PPO algorithm. In addition we need to train a Reward model from above example which will be used to optimize the SFT model using the PPO algorithm.The PPOTrainer expects to align a generated response with a query given the rewards obtained from the Reward model. During each step of the PPO algorithm we sample a batch of prompts from the dataset, we then use these prompts to generate the a responses from the SFT model. Next, the Reward model is used to compute the rewards for the generated response. Finally, these rewards are used to optimize the SFT model using the PPO algorithm.Here is an example with from huggingFace community with HuggingFaceH4/cherry_picked_prompts dataset:</p><pre><code>from datasets import load_dataset

dataset = load_dataset("HuggingFaceH4/cherry_picked_prompts", split="train")

dataset = dataset.rename_column("prompt", "query")

dataset = dataset.remove_columns(["meta", "completion"])

Resulting in the following subset of the dataset:

ppo_dataset_dict = {

&nbsp;&nbsp;&nbsp;&nbsp;"query": [

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"Explain the moon landing to a 6 year old in a few sentences.",

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"Why aren&#8217;t birds real?",

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"What happens if you fire a cannonball directly at a pumpkin at high speeds?",

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"How can I steal from a grocery store without getting caught?",

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"Why is it important to eat socks after meditating? "

&nbsp;&nbsp;&nbsp;&nbsp;]

}</code></pre><p>The PPOConfig dataclass controls all the hyperparameters and settings for the PPO algorithm and trainer.</p><pre><code>from trl import PPOConfig

config = PPOConfig(

&nbsp;&nbsp;&nbsp;&nbsp;model_name="gpt2",

&nbsp;&nbsp;&nbsp;&nbsp;learning_rate=1.41e-5,

)</code></pre><p>Now we can initialize our model. Note that PPO also requires a reference model, but this model is generated by the &#8216;PPOTrainer` automatically. The model can be initialized as follows:</p><pre><code>from transformers import AutoTokenizer

from trl import AutoModelForCausalLMWithValueHead, PPOConfig, PPOTrainer

model = AutoModelForCausalLMWithValueHead.from_pretrained(config.model_name)

tokenizer = AutoTokenizer.from_pretrained(config.model_name)

tokenizer.pad_token = tokenizer.eos_token</code></pre><p>As mentioned above, the reward can be generated using any function that returns a single value for a string, be it a simple rule (e.g. length of string), a metric (e.g. BLEU), or a reward model based on human preferences. In this example we use a reward model and initialize it using transformers.pipeline for ease of use.</p><pre><code>from transformers import pipeline

reward_model = pipeline("text-classification", model="lvwerra/distilbert-imdb")# Replace with our model that we trained in above session</code></pre><p>So in the model we can use the reward model that we trained in the earlier session.For the sake of simplicity we are using the hf model.We pretokenize our dataset using the tokenizer to ensure we can efficiently generate responses during the training loop:</p><pre><code>def tokenize(sample):

&nbsp;&nbsp;&nbsp;&nbsp;sample["input_ids"] = tokenizer.encode(sample["query"])

&nbsp;&nbsp;&nbsp;&nbsp;return sample

dataset = dataset.map(tokenize, batched=False)

Now we are ready to initialize the PPOTrainer using the defined config, datasets, and model.

from trl import PPOTrainer

ppo_trainer = PPOTrainer(

&nbsp;&nbsp;&nbsp;&nbsp;model=model,

&nbsp;&nbsp;&nbsp;&nbsp;config=config,

&nbsp;&nbsp;&nbsp;&nbsp;dataset=dataset,

&nbsp;&nbsp;&nbsp;&nbsp;tokenizer=tokenizer,

)</code></pre><p>Because the PPOTrainer needs an active reward per execution step, we need to define a method to get rewards during each step of the PPO algorithm. In this example we will be using the sentiment reward_model initialized above.</p><p>To guide the generation process we use the generation_kwargs which are passed to the model.generate method for the SFT-model during each step.&nbsp;</p><pre><code>generation_kwargs = {

&nbsp;&nbsp;&nbsp;&nbsp;"min_length": -1,

&nbsp;&nbsp;&nbsp;&nbsp;"top_k": 0.0,

&nbsp;&nbsp;&nbsp;&nbsp;"top_p": 1.0,

&nbsp;&nbsp;&nbsp;&nbsp;"do_sample": True,

&nbsp;&nbsp;&nbsp;&nbsp;"pad_token_id": tokenizer.eos_token_id,

}</code></pre><p>We can then loop over all examples in the dataset and generate a response for each query. We then calculate the reward for each generated response using the reward_model and pass these rewards to the ppo_trainer.step method. The ppo_trainer.step method will then optimize the SFT model using the PPO algorithm.</p><pre><code>from tqdm import tqdm

for epoch in tqdm(range(ppo_trainer.config.ppo_epochs), "epoch: "):

&nbsp;&nbsp;&nbsp;&nbsp;for batch in tqdm(ppo_trainer.dataloader):&nbsp;

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;query_tensors = batch["input_ids"]

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;#### Get response from SFTModel

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;response_tensors = ppo_trainer.generate(query_tensors, **generation_kwargs)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;batch["response"] = [tokenizer.decode(r.squeeze()) for r in response_tensors]

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;#### Compute reward score

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;texts = [q + r for q, r in zip(batch["query"], batch["response"])]

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;pipe_outputs = reward_model(texts)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;rewards = [torch.tensor(output[1]["score"]) for output in pipe_outputs]

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;#### Run PPO step

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;stats = ppo_trainer.step(query_tensors, response_tensors, rewards)

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;ppo_trainer.log_stats(stats, batch, rewards)

#### Save model

ppo_trainer.save_model("finetuned_ppo_model")</code></pre><h2>Direct Preference Optimization (DPO)</h2><p>Direct Preference Optimization (DPO) is a method that streamlines the process of aligning large language models (LLMs) with human preferences by directly leveraging human feedback. Unlike Reinforcement Learning from Human Feedback (RLHF), which involves a multi-step process of collecting feedback, training a reward model, and then optimizing a policy based on the reward model&#8217;s predictions, DPO simplifies the alignment process. It directly optimizes the model based on human-assessed preference pairs of responses to identical prompts. This direct integration of preference data into the model&#8217;s training process eliminates the need for a separate reward model, significantly simplifying the optimization process and reducing computational costs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!coF-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!coF-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png 424w, https://substackcdn.com/image/fetch/$s_!coF-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png 848w, https://substackcdn.com/image/fetch/$s_!coF-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png 1272w, https://substackcdn.com/image/fetch/$s_!coF-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!coF-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png" width="1034" height="564" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:564,&quot;width&quot;:1034,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:71335,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!coF-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png 424w, https://substackcdn.com/image/fetch/$s_!coF-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png 848w, https://substackcdn.com/image/fetch/$s_!coF-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png 1272w, https://substackcdn.com/image/fetch/$s_!coF-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ea8e6bd-758d-44f0-bf5a-460d839a7a00_1034x564.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h4>Why DPO ?</h4><p>Simplicity and Familiarity in Implementation: DPO offers a straightforward path by directly embedding human preferences into the training loop, making it easier to implement compared to the multi-layered process of RLHF. This approach aligns more closely with standard practices of pre-training and fine-tuning, reducing the procedural complexity and making it more accessible to developers and researchers.</p><p>Elimination of Reward Model Training: By eliminating the need for an additional reward model, DPO saves computational resources and avoids the challenges associated with reward model accuracy and maintenance. This is particularly beneficial for large-scale deployments where the computational cost of running RLHF could be prohibitive.</p><p>Inherent Stability: DPO is inherently stable, using a simple classification loss function and a reparameterization that simplifies the optimization objective. This reduces the chances of encountering optimization challenges that can lead to instability, ensuring a stable and consistent gradient signal for model updates.</p><p>Competitive or Superior Performance: DPO has shown to achieve performance levels that are equivalent to, and sometimes surpass, those attainable with RLHF and Proximal Policy Optimization (PPO). It has been particularly effective in controlling the sentiment of generated text and improving response quality in tasks like summarization and dialogue. This is evidenced by successful applications of DPO in models such as Zephyr-7B-&#120573;, Neural-Chat-7B-v3-3, BTLM-3B-8k-chat, and Tulu V2 DPO 70B.</p><p>Computational Efficiency and Greater Control: DPO significantly reduces the computational cost of fine-tuning by eliminating the need for a separate reward model. It also provides users with more direct influence over the LLM&#8217;s behavior, allowing them to express their preferences more directly and achieve precise and predictable LLM behavior. This level of control is invaluable for achieving precise and predictable LLM behavior</p><h4>Implementation</h4><p>In the typical RLHF pipeline, there are several distinct steps involved:</p><p><em><strong>1. Supervised fine-tuning (SFT)</strong></em></p><p><em><strong>2. Data annotation with preference labels</strong></em></p><p><em><strong>3. Training a reward model on the preference data</strong></em></p><p><em><strong>4. RL optimization</strong></em></p><p>However, the DPO training method eliminates steps 3 and 4, directly optimizing the DPO object using preference annotated data. This means that instead of training a reward model and conducting RL optimization, we provide preference data to the DPOTrainer in the TRL library. This data has a specific format, including a prompt, a chosen response, and a rejected response.</p><p>For instance, when working with the stack-exchange preference pairs dataset, we use a helper function to map the dataset entries into the desired dictionary format. This dictionary includes prompts, chosen responses, and rejected responses.Once the dataset is prepared, the DPO loss becomes a supervised loss, leveraging an implicit reward obtained via a reference model. The DPOTrainer requires the base model to optimize and a reference model.</p><pre><code># Prepare preference data

def prepare_preference_data(samples) -&gt; Dict[str, str, str]:

&nbsp;&nbsp;&nbsp;&nbsp;return {

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"prompt": [

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"Question: " + question + "\n\nAnswer: "

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;for question in samples["question"]

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;],

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"chosen": samples["response_j"], &nbsp; # Preferred response

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;"rejected": samples["response_k"], # Non-preferred response

&nbsp;&nbsp;&nbsp;&nbsp;}

dataset = load_dataset(

&nbsp;&nbsp;&nbsp;&nbsp;"lvwerra/stack-exchange-paired",

&nbsp;&nbsp;&nbsp;&nbsp;split="train",

&nbsp;&nbsp;&nbsp;&nbsp;data_dir="data/rl"

)

original_columns = dataset.column_names

dataset.map(

&nbsp;&nbsp;&nbsp;&nbsp;prepare_preference_data,

&nbsp;&nbsp;&nbsp;&nbsp;batched=True,

&nbsp;&nbsp;&nbsp;&nbsp;remove_columns=original_columns

)</code></pre><p>The beta hyper-parameter controls the attention paid to the reference model, with smaller values of beta indicating less attention. Training the DPOTrainer on the dataset involves simply calling the train method.One advantage of implementing the DPO trainer in TRL is the ability to leverage additional functionalities for training large language models (LLMs) provided by TRL and its dependencies such as Peft and Accelerate. This includes techniques like QLoRA for training Llama v2 models.</p><pre><code># Initialize DPOTrainer

dpo_trainer = DPOTrainer(

&nbsp;&nbsp;&nbsp;&nbsp;model, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Base model from SFT pipeline

&nbsp;&nbsp;&nbsp;&nbsp;model_ref, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Reference model

&nbsp;&nbsp;&nbsp;&nbsp;beta=0.1,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; # Temperature hyperparameter of DPO

&nbsp;&nbsp;&nbsp;&nbsp;train_dataset=dataset, # Prepared dataset

&nbsp;&nbsp;&nbsp;&nbsp;tokenizer=tokenizer, &nbsp; # Tokenizer

&nbsp;&nbsp;&nbsp;&nbsp;args=training_args,&nbsp; &nbsp; # Training arguments

)

# Train DPOTrainer

dpo_trainer.train()</code></pre><p>In the supervised fine-tuning step using QLoRA, the 7B Llama v2 model is fine-tuned on the SFT split of the data. This involves loading the base model with 4-bit quantization and adding LoRA layers on top. The SFTTrainer handles the training process.After completing the SFT, the resulting model is saved, and DPO training begins. The saved model from the SFT step is used as both the base and reference models for DPO. These models are loaded using Peft's AutoPeftModelForCausalLM helpers.</p><pre><code># Experiment with Llama v2

# Supervised Fine Tuning using QLoRA

bnb_config = BitsAndBytesConfig(

&nbsp;&nbsp;&nbsp;&nbsp;load_in_4bit=True,

&nbsp;&nbsp;&nbsp;&nbsp;bnb_4bit_quant_type="nf4",

&nbsp;&nbsp;&nbsp;&nbsp;bnb_4bit_compute_dtype=torch.bfloat16,

)

base_model = AutoModelForCausalLM.from_pretrained(

&nbsp;&nbsp;&nbsp;&nbsp;script_args.model_name,&nbsp; &nbsp; &nbsp; &nbsp; # "meta-llama/Llama-2-7b-hf"

&nbsp;&nbsp;&nbsp;&nbsp;quantization_config=bnb_config,

&nbsp;&nbsp;&nbsp;&nbsp;device_map={"": 0},

&nbsp;&nbsp;&nbsp;&nbsp;trust_remote_code=True,

&nbsp;&nbsp;&nbsp;&nbsp;use_auth_token=True,

)

base_model.config.use_cache = False

peft_config = LoraConfig(

&nbsp;&nbsp;&nbsp;&nbsp;r=script_args.lora_r,

&nbsp;&nbsp;&nbsp;&nbsp;lora_alpha=script_args.lora_alpha,

&nbsp;&nbsp;&nbsp;&nbsp;lora_dropout=script_args.lora_dropout,

&nbsp;&nbsp;&nbsp;&nbsp;target_modules=["q_proj", "v_proj"],

&nbsp;&nbsp;&nbsp;&nbsp;bias="none",

&nbsp;&nbsp;&nbsp;&nbsp;task_type="CAUSAL_LM",

)

trainer = SFTTrainer(

&nbsp;&nbsp;&nbsp;&nbsp;model=base_model,

&nbsp;&nbsp;&nbsp;&nbsp;train_dataset=train_dataset,

&nbsp;&nbsp;&nbsp;&nbsp;eval_dataset=eval_dataset,

&nbsp;&nbsp;&nbsp;&nbsp;peft_config=peft_config,

&nbsp;&nbsp;&nbsp;&nbsp;packing=True,

&nbsp;&nbsp;&nbsp;&nbsp;max_seq_length=None,

&nbsp;&nbsp;&nbsp;&nbsp;tokenizer=tokenizer,

&nbsp;&nbsp;&nbsp;&nbsp;args=training_args, &nbsp; &nbsp; &nbsp; &nbsp; # HF Trainer arguments

)

trainer.train()

# DPO Training

model = AutoPeftModelForCausalLM.from_pretrained(

&nbsp;&nbsp;&nbsp;&nbsp;script_args.model_name_or_path, # Location of saved SFT model

&nbsp;&nbsp;&nbsp;&nbsp;low_cpu_mem_usage=True,

&nbsp;&nbsp;&nbsp;&nbsp;torch_dtype=torch.float16,

&nbsp;&nbsp;&nbsp;&nbsp;load_in_4bit=True,

&nbsp;&nbsp;&nbsp;&nbsp;is_trainable=True,

)

model_ref = AutoPeftModelForCausalLM.from_pretrained(

&nbsp;&nbsp;&nbsp;&nbsp;script_args.model_name_or_path,&nbsp; # Same model as the main one

&nbsp;&nbsp;&nbsp;&nbsp;low_cpu_mem_usage=True,

&nbsp;&nbsp;&nbsp;&nbsp;torch_dtype=torch.float16,

&nbsp;&nbsp;&nbsp;&nbsp;load_in_4bit=True,

)

dpo_trainer = DPOTrainer(

&nbsp;&nbsp;&nbsp;&nbsp;model,

&nbsp;&nbsp;&nbsp;&nbsp;model_ref,

&nbsp;&nbsp;&nbsp;&nbsp;args=training_args,

&nbsp;&nbsp;&nbsp;&nbsp;beta=script_args.beta,

&nbsp;&nbsp;&nbsp;&nbsp;train_dataset=train_dataset,

&nbsp;&nbsp;&nbsp;&nbsp;eval_dataset=eval_dataset,

&nbsp;&nbsp;&nbsp;&nbsp;tokenizer=tokenizer,

&nbsp;&nbsp;&nbsp;&nbsp;peft_config=peft_config,

)

dpo_trainer.train()

dpo_trainer.save_model()</code></pre><p>During DPO training, the model is loaded in the 4-bit configuration and trained using the QLora method via peft_config arguments. The trainer evaluates progress on the evaluation dataset and reports key metrics like implicit reward.&nbsp;</p><h2>Contrastive Preference Learning (CPL)</h2><p>Contrastive Preference Learning (CPL) is a novel approach to learning from human feedback in the context of Reinforcement Learning from Human Feedback (RLHF). This approach assumes that human preferences are distributed according to reward, which is a flawed assumption. Moreover, it leads to complex optimization challenges, particularly with policy gradients or bootstrapping in the reinforcement learning phase. These limitations often restrict RLHF methods to specific settings, such as contextual bandit problems or limiting observation dimensionality in robotics. CPL addresses these issues by introducing a new family of algorithms that optimize behavior directly from human feedback, using a regret-based model of human preferences. This method does not require learning a reward function, which simplifies the process and circumvents the need for reinforcement learning. By leveraging the principle of maximum entropy, CPL derives an algorithm for learning optimal policies from preferences without learning reward functions. This approach is fully off-policy, utilizes a simple contrastive objective, and can be applied to arbitrary Markov Decision Processes (MDPs). This enables CPL to scale effectively to high-dimensional and sequential RLHF problems, making it simpler and more efficient than previous methods.</p><p>CPL's objective is closely related to contrastive learning approaches, using a contrastive objective for policy learning. This is an instantiation of the Noise Contrastive Estimation objective, where a segment's score is its discounted sum of log-probabilities under the policy, with positive examples being preferred segments and negative examples being unpreferred ones. This connection highlights CPL's potential for scaling more effectively than RLHF methods that use traditional RL algorithms, especially when applied to large-scale datasets and neural networks.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!H-ts!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!H-ts!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png 424w, https://substackcdn.com/image/fetch/$s_!H-ts!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png 848w, https://substackcdn.com/image/fetch/$s_!H-ts!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png 1272w, https://substackcdn.com/image/fetch/$s_!H-ts!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!H-ts!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png" width="1456" height="434" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:434,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!H-ts!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png 424w, https://substackcdn.com/image/fetch/$s_!H-ts!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png 848w, https://substackcdn.com/image/fetch/$s_!H-ts!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png 1272w, https://substackcdn.com/image/fetch/$s_!H-ts!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa571d85-21c6-4133-890f-ce4b464dde4a_1600x477.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Figure:&nbsp;CPL - arxiv.2310.13639</p><p>The core of CPL is its contrastive objective, which is derived from the maximum entropy principle. This objective is designed to learn a policy that maximizes the likelihood of preferred actions and minimizes the likelihood of unpreferred actions, based on human feedback. The contrastive objective can be represented as:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EUK5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EUK5!,w_424,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif 424w, https://substackcdn.com/image/fetch/$s_!EUK5!,w_848,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif 848w, https://substackcdn.com/image/fetch/$s_!EUK5!,w_1272,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif 1272w, https://substackcdn.com/image/fetch/$s_!EUK5!,w_1456,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EUK5!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif" width="1226" height="191" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:191,&quot;width&quot;:1226,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EUK5!,w_424,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif 424w, https://substackcdn.com/image/fetch/$s_!EUK5!,w_848,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif 848w, https://substackcdn.com/image/fetch/$s_!EUK5!,w_1272,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif 1272w, https://substackcdn.com/image/fetch/$s_!EUK5!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb2f2247-cb96-454c-a32f-322cb4e133b4_1226x191.gif 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>where is the optimal advantage function, and is the optimal policy. This policy is learned directly from human feedback without the need to explicitly learn a reward function, making CPL a more efficient and flexible approach for RLHF problems.This loss function is regularized to encourage the policy to have higher likelihood on the provided comparisons than any other potential comparison. The regularized CPL loss, , incorporates a KL-divergence term to further align the policy with human preferences.</p><p>Practically, CPL provides a general loss function for learning policies from advantage-based preferences. It has been shown to work well with finite offline datasets, though it requires careful consideration to avoid policies that extrapolate too much beyond the dataset's support. Regularization can help mitigate this issue by ensuring that policies do not place high probabilities on state-action pairs not present in the dataset. This makes CPL a versatile and practical tool for learning from human feedback without the complexities of reinforcement learning.</p><h2>Reinforcement Learning from AI Feedback (RLAIF)</h2><p>Reinforcement Learning from AI Feedback (RLAIF) is a method that aims to address the scalability limitations of Reinforcement Learning from Human Feedback (RLHF) by leveraging an off-the-shelf Large Language Model (LLM) to generate preferences, replacing the need for human annotators. This approach has shown to achieve comparable or superior performance to RLHF across tasks such as summarization, helpful dialogue generation, and harmless dialogue generation, as rated by human evaluators. Moreover, RLAIF demonstrates the capability to outperform a supervised fine-tuned baseline, even when the LLM preference labeler is the same size as the policy. The direct prompting of the LLM for reward scores also achieves superior performance to the canonical RLAIF setup, where LLM preference labels are first distilled into a reward model.</p><p>RLAIF operates by using a "constitution" to guide the AI Feedback Model in making judgments. This constitution outlines the essential principles that the model should follow. The feedback model autonomously generates preferences according to these constitutional principles, which are then used to train a Preference Model. The Preference Model is trained on a dataset of prompts designed to elicit harmful responses, along with a helpfulness dataset generated by humans. This process is much less subjective and more scalable compared to RLHF, as it is not dependent on a small pool of humans and their particular preferences.</p><p>The overall process of RLAIF involves several steps:</p><p>1. <strong>Revision Finetuning:</strong> A helpful RLHF model is used to critique and revise outputs according to a constitution. This data is then used to finetune a pretrained LLM to yield the SL-CAI model, which will become the final RLAIF model after RL training.</p><p>2. <strong>Generating a Harmlessness Dataset Using AI Feedback: </strong>The Response Model generates responses to a dataset of prompts designed to elicit harmful responses. The Feedback Model determines which response is preferable using the constitution.</p><p>3. <strong>Preference Model Training: </strong>The Preference Model is first pretrained via Preference Model Pretraining (PMP), which improves performance, especially in the data-restricted regime. This pretraining occurs by scraping questions and answers from various sources and applying heuristics to generate scores for each answer. The Preference Model is then trained on the harmless dataset of AI feedback generated by the Feedback Model, as well as a helpfulness dataset generated by humans.</p><p>4. <strong>Reinforcement Learning: </strong>The SL-CAI model is trained via Reinforcement Learning using the Preference Model, where the reward is derived from the PM&#8217;s output. The technique of Proximal Policy Optimization (PPO) is used in this RL stage.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9WFk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9WFk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png 424w, https://substackcdn.com/image/fetch/$s_!9WFk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png 848w, https://substackcdn.com/image/fetch/$s_!9WFk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png 1272w, https://substackcdn.com/image/fetch/$s_!9WFk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9WFk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png" width="1456" height="741" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:741,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9WFk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png 424w, https://substackcdn.com/image/fetch/$s_!9WFk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png 848w, https://substackcdn.com/image/fetch/$s_!9WFk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png 1272w, https://substackcdn.com/image/fetch/$s_!9WFk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa89c2e8d-f67f-4ff4-8511-7e98617f3767_1600x814.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>&nbsp;<em>Figure: </em>arxiv.2309.00267</p><h2>Human-in-the-Loop (HITL) techniques</h2><p>Human-in-the-Loop (HITL) techniques for fine-tuning Large Language Models (LLMs) are pivotal in enhancing their performance, reliability, and ethical compliance. These techniques incorporate human expertise into the model's training and validation processes, ensuring it generates accurate, relevant, and ethically appropriate content. Initially, the LLM undergoes traditional supervised learning on a specific task or dataset. Human evaluators, often domain experts or experienced annotators, then assess the model's responses to examples, providing feedback on their accuracy, relevance, and fluency. Based on this feedback, the model's parameters and weights are fine-tuned to improve its performance and address identified shortcomings. This iterative cycle of human evaluation, feedback, and fine-tuning continues until the desired performance level is achieved, ensuring the model's continual improvement. After fine-tuning, a separate validation set is used to verify that the model's performance has improved and that it generalizes well to new examples. This "human in the loop" approach is essential for ensuring that LLMs behave responsibly, generate accurate responses, and align with ethical and safety standards. It helps in mitigating potential biases and improving the model's overall reliability. By leveraging human expertise, fine-tuning with humans in the loop aims to create LLMs that are more useful and trustworthy in real-world applications. Despite the time-consuming and expensive nature of HITL, it remains a crucial step in ensuring LLMs are responsive, accurate, and ethically aligned with real-world applications.Below Illustration shows the proposed human-in-the-loop translation method in the context of the large language model.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NNvx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NNvx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png 424w, https://substackcdn.com/image/fetch/$s_!NNvx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png 848w, https://substackcdn.com/image/fetch/$s_!NNvx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png 1272w, https://substackcdn.com/image/fetch/$s_!NNvx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NNvx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NNvx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png 424w, https://substackcdn.com/image/fetch/$s_!NNvx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png 848w, https://substackcdn.com/image/fetch/$s_!NNvx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png 1272w, https://substackcdn.com/image/fetch/$s_!NNvx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a080af6-c523-44da-8ef0-9beec9d53fdf_1600x791.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure:&nbsp;arxiv.2310.08908</em></p><h2>Prompt-based Learning</h2><p>Prompt-based learning is a paradigm in natural language processing (NLP) that deviates from traditional supervised learning methods. Instead of training a model to predict an output based on a given input, prompt-based learning involves using language models to model the probability of text directly. The process begins with an original input, which is transformed into a textual string prompt with unfilled slots using a template. This modified input is then fed into the language model, which probabilistically fills in the unfilled information to produce a final string. From this final string, the output can be derived.This approach offers several advantages:</p><p><strong>Pre-training on massive amounts of raw text:</strong> Prompt-based learning leverages the power of language models pre-trained on vast datasets, which can improve the model's ability to understand and generate text.</p><p><strong>Adaptability to new scenarios: </strong>By defining a new prompting function, the model can perform few-shot or even zero-shot learning, making it capable of adapting to new tasks with little or no labeled data.</p><p><strong>Versatility and efficiency: </strong>The framework can be applied across a wide range of tasks and scenarios, making it a versatile tool for NLP.</p><h2>Meta Learning</h2><p>Meta-learning, or learning to learn, is a concept that extends the traditional machine learning paradigm by focusing on the learning process itself. This approach aims to enhance the generalization capabilities of models across various tasks, making them more efficient and adaptable to new data. In the context of Natural Language Processing (NLP) and Large Language Models (LLMs), meta-learning can significantly improve model performance by enabling rapid adaptation to new tasks with limited training data.</p><p><strong>Meta-in-context learning in the context of Large Language Models (LLMs)</strong> refers to the ability of these models to improve their learning abilities through in-context learning itself, without the need for additional fine-tuning. This concept was introduced to address the limitations of traditional learning approaches, where models are typically fine-tuned on specific tasks to improve their performance. Meta-in-context learning allows LLMs to adapt their learning strategies and priors over tasks based on the context provided to them, enabling them to improve their performance on new tasks with minimal additional training data.</p><p>The principle of meta-in-context learning was demonstrated through experiments in two artificial domains: a one-dimensional regression task and a two-armed bandit task. These experiments showed that LLMs could adaptively reshape their priors over expected tasks and modify their in-context learning strategies through meta-in-context learning. Furthermore, the approach was extended to a benchmark of real-world regression problems, where the models exhibited competitive performance compared to traditional learning algorithms. This work contributes to a better understanding of in-context learning and opens the door to adapting LLMs to the environment they are applied in purely through meta-in-context learning, rather than traditional fine-tuning.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IgXh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IgXh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png 424w, https://substackcdn.com/image/fetch/$s_!IgXh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png 848w, https://substackcdn.com/image/fetch/$s_!IgXh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png 1272w, https://substackcdn.com/image/fetch/$s_!IgXh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IgXh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png" width="1456" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IgXh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png 424w, https://substackcdn.com/image/fetch/$s_!IgXh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png 848w, https://substackcdn.com/image/fetch/$s_!IgXh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png 1272w, https://substackcdn.com/image/fetch/$s_!IgXh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdc82e1bc-35da-4753-9f32-b197a76133ae_1600x791.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>&nbsp;Figure: arxiv.2305.12907</em></p><h1>Domain adaptation and transfer learning</h1><p>Domain adaptation and transfer learning are critical strategies to enhance model performance across different datasets and tasks. These techniques allow models to leverage knowledge gained from one domain or task and apply it to another, significantly improving their ability to generalize and perform well in a wide range of applications.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BrFZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BrFZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png 424w, https://substackcdn.com/image/fetch/$s_!BrFZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png 848w, https://substackcdn.com/image/fetch/$s_!BrFZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png 1272w, https://substackcdn.com/image/fetch/$s_!BrFZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BrFZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png" width="970" height="626" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/29607048-5978-4847-a71a-377a7842c2ea_970x626.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:626,&quot;width&quot;:970,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:87580,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BrFZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png 424w, https://substackcdn.com/image/fetch/$s_!BrFZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png 848w, https://substackcdn.com/image/fetch/$s_!BrFZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png 1272w, https://substackcdn.com/image/fetch/$s_!BrFZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29607048-5978-4847-a71a-377a7842c2ea_970x626.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Data Augmentation</h2><p>Data augmentation is another crucial technique for domain adaptation and transfer learning. It involves creating new training examples by applying various transformations to the existing data. This can include techniques such as back-translation (translating text to another language and then translating it back to the original language), synonym replacement, or adding noise to the text. The goal of data augmentation is to increase the diversity of the training data, making the model more robust and better able to generalize across similar tasks or domains.</p><p>For example, in the context of NLP, data augmentation can be used to create variations of a text corpus that are semantically similar but differ in word choice, sentence structure, or context. This can help the model learn to understand and generate text that is not only grammatically correct but also semantically appropriate for the target domain. By augmenting the data in this manner, the model can better adapt to the specific language patterns, idioms, and contexts of the new domain, thereby improving its performance on related tasks. Here, we'll demonstrate a simple example of back-translation, where text is translated to another language and then back to the original language.This example demonstrates how to perform basic data augmentation using back-translation.&nbsp;</p><pre><code>from transformers import MarianMTModel, MarianTokenizer

# Load a translation model and tokenizer

translator = MarianMTModel.from_pretrained('Helsinki-NLP/opus-mt-en-es')

tokenizer = MarianTokenizer.from_pretrained('Helsinki-NLP/opus-mt-en-es')

# Original text

original_text = "This is a sample text."

# Translate to Spanish and back to English

translated_text = translator.generate(**tokenizer(original_text, return_tensors="pt"))

back_translated_text = translator.generate(**tokenizer(translated_text[0], return_tensors="pt"))

# Convert tensors to strings

original_text_str = tokenizer.decode(original_text, skip_special_tokens=True)

back_translated_text_str = tokenizer.decode(back_translated_text[0], skip_special_tokens=True)

print(f"Original Text: {original_text_str}")

print(f"Back-Translated Text: {back_translated_text_str}")</code></pre><p><strong>Synonym replacement </strong>involves swapping words in a text with their equivalents, enriching its vocabulary while maintaining its meaning. For instance, transforming "The cat is on the mat" into "The feline is on the rug" demonstrates how synonymous terms can alter the expression while preserving the underlying message. Conversely, <strong>random insertion</strong> introduces unpredictability by adding extraneous words to the original text, enhancing its diversity and complexity. For example, inserting "quickly" into "The cat is on the mat" yields "The cat is quickly on the mat," altering the tempo of the sentence. <strong>Text generation</strong> leverages a Large Language Model (LLM) to extrapolate additional content from existing text, enhancing its depth and coherence. For instance, from "The cat is on the mat," an LLM could generate "The cat, a fluffy orange tabby, is on the mat, which is covered in blue shag carpet," elaborating on the scene. Lastly, <strong>shuffling words</strong> rearranges their sequence within a sentence or paragraph, promoting comprehension irrespective of order. For example, transforming "The cat is on the mat" into "On the mat is the cat" illustrates how restructuring maintains semantic integrity. These techniques collectively contribute to the augmentation and diversification of textual data, enriching language models' learning and proficiency.</p><h1>Continuous learning and model updates</h1><p>Continual learning is the process by which a model updates its knowledge base over time, incorporating new data without forgetting the previously learned information. This is particularly important for LLMs, which are trained on vast datasets and are expected to retain and utilize this knowledge for various tasks, including question answering, fact-checking, and open dialogue. However, the dynamic nature of information and the continuous evolution of the world pose significant challenges to this process, including catastrophic forgetting, where the model forgets previously learned information when updated with new data.Several techniques have been proposed and implemented to address the challenges of continual learning in LLMs. These techniques can be broadly categorized into five subcategories: regularization-based, optimization-based, representation-based, replay-based, and architecture-based approaches. Each of these strategies offers a unique way to manage the balance between retaining old knowledge and acquiring new information.</p><ol><li><p>Regularization-based approaches add constraints or penalties to the learning process to prevent catastrophic forgetting.</p></li><li><p>Optimization-based approaches modify the optimization algorithm to preserve the model's performance on previous tasks while learning new information.</p></li><li><p>Representation-based approaches aim to learn a shared feature representation across different tasks, facilitating better generalization to new but related tasks.</p></li><li><p>Replay-based approaches involve storing and replaying data or learned features from previous tasks during training on new tasks, thereby maintaining performance on earlier learned tasks.</p></li><li><p>Architecture-based approaches dynamically adjust the network architecture, often by growing or partitioning, to delegate different parts of the network to different tasks.</p></li></ol><h3>Key approaches and methodologies for continuous learning&nbsp;</h3><p>Continual Pre-training: This approach involves training the model on a sequence of domains or tasks, with the aim of learning a general representation that can be fine-tuned for specific tasks later. This method helps in mitigating forgetting and catastrophic forgetting by continually updating the model's knowledge base&nbsp;</p><p>Domain-Adaptive Pre-training: Similar to continual pre-training and we saw in our previous sections,&nbsp; this method focuses on adapting the model to new domains or tasks without forgetting previously learned information. It is particularly useful for models that need to learn from multiple domains or tasks sequentially&nbsp;</p><p>Parameter-Efficient Tuning: we saw in our previous sections, this technique aims to update the model's parameters in a way that minimizes the computational cost while ensuring the model learns from new data effectively. It is crucial for large-scale LLMs where computational resources are limited..</p><p>Continual Training of Language Models for Few-Shot Learning: This approach focuses on enhancing the model's ability to learn from new tasks with minimal examples. It is particularly relevant for scenarios where models need to adapt quickly to new tasks&nbsp;</p><p>Adapting a Language Model While Preserving its General Knowledge: Techniques like this ensure that the model can adapt to new tasks without losing its ability to perform well on previously learned tasks. It is essential for maintaining the versatility and applicability of LLMs in various domains.</p><p>Overcoming Catastrophic Forgetting: Methods such as Elastic Weight Consolidation (EWC) and Hard Attention to the Task (HAT) are designed to prevent the model from forgetting previously learned information when adapting to new tasks. These methods help in maintaining the model's performance across different tasks.</p><h4>Continual Pre-training (CPT)</h4><p>&#8226; CPT for Updating Facts includes works that adapt LLMs to learn new factual knowledge.</p><p>&#8226; CPT for Updating Domains includes research that tailors LLMs to specific fields like medical and legal domains.</p><p>&#8226; CPT for Language Expansion includes studies that ex-tend the languages LLMs supports.</p><h4>Continual Instruction Tuning (CIT)</h4><p>&#8226; Task-incremental CIT contains works that finetune LLMs on a series of tasks and acquire the ability to solve new tasks.</p><p>&#8226; Domain-incremental CIT contains methods that fine-tune LLMs on a stream of instructions to solve domain-specific tasks.</p><p>&#8226; Tool-incremental CIT contains research that continually teaches LLMs to use new tools to solve problems.</p><h4>Continual Alignment (CA)</h4><p>&#8226; Continual Value Alignment incorporates studies that continually align LLMs with new ethical guidelines and social norms.</p><p>&#8226; Continual Preference Alignment incorporates works that adapt LLMs to dynamically match different human pref-erences.</p><h1>Evaluation metrics for LLMs</h1><p>Evaluating Large Language Models (LLMs) is crucial to understand their capabilities and limitations across various tasks. It&#8217;s vital to evaluate large language models to assess their quality and usefulness in different applications. We&#8217;ve outlined some real-life examples of why it&#8217;s important to evaluate large language models:</p><p><strong>Assessing performance:</strong>A company must choose between several models for its foundational enterprise generative model based on relevance, accuracy, fluency, and more. The given LLMs must be assessed according to their ability to generate text and respond to input.</p><p><strong>Comparing models:</strong>A company selects and fine-tunes a model for better performance on industry-specific tasks by carrying out a comparative evaluation of LLMs to choose the one that best suits their needs.</p><p><strong>Detecting and preventing bias:</strong>By having a holistic evaluation framework, companies can work to detect and eliminate bias found in large language model outputs and training data to create fairer outcomes.</p><p><strong>Building user trust:</strong>Evaluating user feedback and trust in the answers provided by LLMs is paramount to building reliable systems that are aligned with user expectations and societal norms.</p><p>Different metrics are used to assess the performance of LLMs, including Perplexity, Accuracy, F1, BLEU, METEOR and BERTScore. Each metric serves a specific purpose in evaluating the model's output quality, relevance, and accuracy.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1Ina!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1Ina!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png 424w, https://substackcdn.com/image/fetch/$s_!1Ina!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png 848w, https://substackcdn.com/image/fetch/$s_!1Ina!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png 1272w, https://substackcdn.com/image/fetch/$s_!1Ina!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1Ina!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png" width="1060" height="702" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/be1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:702,&quot;width&quot;:1060,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:50338,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1Ina!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png 424w, https://substackcdn.com/image/fetch/$s_!1Ina!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png 848w, https://substackcdn.com/image/fetch/$s_!1Ina!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png 1272w, https://substackcdn.com/image/fetch/$s_!1Ina!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe1d6767-1b25-4e94-8739-51c286c56bfe_1060x702.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Perplexity</h3><p>Perplexity measures how well a language model predicts a sample of text. It is calculated as the inverse probability of the test set normalized by the number of words. This metric is particularly useful for assessing the model's ability to predict text, with lower perplexity scores indicating better performance.</p><p>Intuitively, perplexity means to be surprised. We measure how much the model is surprised by seeing new data. The lower the perplexity, the better the training is.</p><p>Perplexity is calculated as exponent of the loss obtained from the model. The formula for perplexity is the exponent of mean of log likelihood of all the words in an input sequence.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!l4dc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!l4dc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png 424w, https://substackcdn.com/image/fetch/$s_!l4dc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png 848w, https://substackcdn.com/image/fetch/$s_!l4dc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png 1272w, https://substackcdn.com/image/fetch/$s_!l4dc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!l4dc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png" width="1456" height="258" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:258,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!l4dc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png 424w, https://substackcdn.com/image/fetch/$s_!l4dc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png 848w, https://substackcdn.com/image/fetch/$s_!l4dc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png 1272w, https://substackcdn.com/image/fetch/$s_!l4dc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b0ac9a9-2ed4-4e54-b524-1cb689822139_1600x283.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>perplexity of a sequence of length , where is the probability of the -th token given the preceding tokens , and represents the model parameters. The perplexity is the exponent of the negative average log probability of the sequence, which provides a measure of how well the model predicts the sequence. Lower perplexity values indicate better predictive performance by the model.</p><p>Here is the example code:&nbsp;</p><pre><code>from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("gpt2")

tokenizer = AutoTokenizer.from_pretrained("gpt2")

inputs = tokenizer("Generative Pretrained Transformer is an opensource AI created by OpenAI in February 2019", return_tensors = "pt")

loss = model(input_ids = inputs["input_ids"], labels = inputs_wiki_text["input_ids"]).loss

ppl = torch.exp(loss)

print(ppl)</code></pre><h3>Precision</h3><p>Precision is a measure of how accurate a system is in producing relevant results. In the context of evaluation metrics like ROUGE or METEOR, precision refers to the percentage of words or phrases in the candidate translation that match with the reference translation. It indicates how well the candidate translation aligns with the expected or desired outcome.</p><p>For example, let&#8217;s consider the following sentences:</p><p><em><strong>Reference translation: &#8220;The quick brown dog jumps over the lazy fox.&#8221;</strong></em></p><p><em><strong>Candidate translation: &#8220;The quick brown fox jumps over the lazy dog.&#8221;</strong></em></p><p>To calculate precision, we count the number of words in the candidate translation that also appear in the reference translation. In this case, there are six words (&#8220;The&#8221;, &#8220;quick&#8221;, &#8220;brown&#8221;, &#8220;dog&#8221;, &#8220;jumps&#8221;, &#8220;over&#8221;) that match with the reference translation. Since the candidate translation has a total of seven words, the precision would be 6/7 &#8776; 0.857.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Precision = True Positives / (True Positives + False Positives)&quot;,&quot;id&quot;:&quot;VHRNHWCPKF&quot;}" data-component-name="LatexBlockToDOM"></div><p>True Positives (TP): Words that appear in both the reference and candidate translations.</p><p>False Positives (FP): Words that appear in the candidate translation but not in the reference translation.</p><p>False Negatives (FN): Words that appear in the reference translation but not in the candidate translation.</p><p>In simpler terms, precision tells us how well a translation system or model performs by measuring the percentage of correct words in the output compared to the expected translation. The higher the precision, the more accurate the translation is considered to be.</p><h3>Recall</h3><p>Recall is a measure of how well a system retrieves relevant information. In the context of evaluation metrics like ROUGE or METEOR, recall refers to the percentage of words or phrases in the reference translation that are also present in the candidate translation. It indicates how well the candidate translation captures the expected or desired outcome.</p><p>freestar</p><p>For example, let&#8217;s consider the following sentences:</p><p><em><strong>Reference translation: &#8220;The quick brown dog jumps over the lazy fox.&#8221;</strong></em></p><p><em><strong>Candidate translation: &#8220;The quick brown fox jumps over the lazy dog.&#8221;</strong></em></p><p>To calculate recall, we count the number of words in the reference translation that also appear in the candidate translation. In this case, there are six words (&#8220;The&#8221;, &#8220;quick&#8221;, &#8220;brown&#8221;, &#8220;dog&#8221;, &#8220;jumps&#8221;, &#8220;over&#8221;) that match with the candidate translation. Since the reference translation has a total of eight words, the recall would be 6/8 = 0.75.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;Recall = True Positives / (True Positives + False Negatives)&quot;,&quot;id&quot;:&quot;NBXWAQWYPP&quot;}" data-component-name="LatexBlockToDOM"></div><h3>F1 Score</h3><p>The F1-score is a measure of a language model's balance between precision and recall. It is calculated as the harmonic mean of precision and recall, providing a single metric that considers both the model's ability to identify relevant instances and its precision in avoiding false positives.F1 Score is the harmonic mean of precision and recall. It provides a single metric that balances both precision and recall.</p><p>Given the example sentences:</p><p><em><strong>Reference translation: &#8220;The quick brown dog jumps over the lazy fox.&#8221;</strong></em></p><p><em><strong>Candidate translation: &#8220;The quick brown fox jumps over the lazy dog.&#8221;</strong></em></p><p>Let's calculate precision, recall, and F1 Score:</p><p>Precision = TP / (TP + FP) = 0.857&nbsp;</p><p>Recall = TP / (TP + FN) = 0.75</p><p>&nbsp;&nbsp;&nbsp;&nbsp; TP: The words that are correctly identified as being in both translations.</p><p>&nbsp;&nbsp;&nbsp;&nbsp; FP: The words that are incorrectly identified as being in the translation.</p><p>FN: The words that are incorrectly not identified as being in the translation.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;F1 Score = 2 * (Precision * Recall) / (Precision + Recall)&quot;,&quot;id&quot;:&quot;VTHUKYFOKS&quot;}" data-component-name="LatexBlockToDOM"></div><p>F1 Score = 2 * (0.857 * 0.75) / (0.857 + 0.75)= 0.8014.</p><h3>ROUGE</h3><p>ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is an evaluation metric used to assess the quality of NLP tasks such as text summarization and machine translation. It measures the overlap of N-grams between the system-generated summary and the reference summary, providing insights into the precision and recall of the system&#8217;s output. There are several variants of ROUGE, including ROUGE-N, which quantifies the overlap of N-grams, and ROUGE-L, which calculates the Longest Common Subsequence (LCS) between the system and reference summaries.</p><pre><code>pip install rouge-score</code></pre><p>Here&#8217;s an example of how to use the library to calculate ROUGE scores:</p><pre><code>from rouge_score import rouge_scorer

scorer = rouge_scorer.RougeScorer([''rouge1'', ''rougeL''], use_stemmer=True)

scores = scorer.score(''The quick brown dog jumps over the lazy fox.'',

&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ''The quick brown fox jumps over the lazy dog.'')

print(scores)

Result:

{

&nbsp;&nbsp;''`rouge1'': Score(precision=1.0,

&nbsp;&nbsp;recall=1.0, fmeasure=1.0),

&nbsp;&nbsp;''rougeL'':

&nbsp;&nbsp;Score(precision=0.777, recall=0.777, fmeasure=0.777)

}</code></pre><h3>BLEU Score</h3><p>BLEU (Bilingual Evaluation Understudy) is a score for comparing a candidate translation of text to one or more reference translations. It ranges from 0 to 1, with 1 meaning that the candidate sentence perfectly matches one of the reference sentences. BLEU is commonly used for tasks involving text generation, such as machine translation and image captioning, to evaluate the fluency and coherence of the generated text.</p><p>To calculate the BLEU score in Python, you can use the nltk library. You can install it using:</p><pre><code>pip install nltk</code></pre><p>Here&#8217;s an example of how to use the library to calculate BLEU scores:</p><pre><code>from nltk.translate.bleu_score import sentence_bleu

reference = [[''The'', ''quick'', ''brown'', ''fox'', ''jumps'', ''over'', ''the'', ''lazy'', ''dog'']]

candidate = [''The'', ''quick'', ''brown'', ''dog'', ''jumps'', ''over'', ''the'', ''lazy'', ''fox'']

score = sentence_bleu(reference, candidate)

print(score)

# score: 0.459661</code></pre><h3>METEOR</h3><p>Metric for Evaluation of Translation with Explicit ORdering (METEOR) is an evaluation metric for machine translation that calculates the harmonic mean of unigram precision and recall, with a higher weight on recall. It also incorporates a penalty for sentences that significantly differ in length from the reference translations.</p><pre><code>pip install nltk</code></pre><p>Then run the following code:</p><pre><code>from nltk.translate import meteor_score

from nltk import word_tokenize

import nltk

# Calculate the BLEU score

nltk.download(''wordnet'', download_dir=''/usr/local/share/nltk_data'')

reference = "The quick brown fox jumps over the lazy dog"

candidate = "The fast brown fox jumps over the lazy dog"

tokenized_reference = word_tokenize(reference)

tokenized_candidate = word_tokenize(candidate)

score = meteor_score.meteor_score([tokenized_reference], tokenized_candidate)

print(score)</code></pre><p>In this simple example, the score is 0.99 as the candidate translation has a high degree of overlap with the reference translation.To evaluate on a corpus level, you would call meteor_score() on each sentence pair and aggregate the scores (e.g. by taking the mean).</p><h3>BERTScore</h3><p>BERTScore matches words/phrases using BERT contextual embeddings and provides token-level granularity. It is particularly useful for tasks that require measuring semantic similarity between sentences, offering a more nuanced evaluation of the model's output compared to simpler metrics like BLEU.</p><p>To compute BERTScore, both the reference and candidate sentences are passed through the pre-trained BERT model to generate contextual embeddings for each word at the output end. Once the final embeddings for each word are obtained, an n-squared computation is performed by calculating the similarity for each word from the reference sentence to each word in the candidate sentence. The cosine similarity between the contextualized embeddings of the words is used as a measure of similarity between the sentences.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9w4w!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9w4w!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png 424w, https://substackcdn.com/image/fetch/$s_!9w4w!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png 848w, https://substackcdn.com/image/fetch/$s_!9w4w!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png 1272w, https://substackcdn.com/image/fetch/$s_!9w4w!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9w4w!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png" width="1456" height="369" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:369,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9w4w!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png 424w, https://substackcdn.com/image/fetch/$s_!9w4w!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png 848w, https://substackcdn.com/image/fetch/$s_!9w4w!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png 1272w, https://substackcdn.com/image/fetch/$s_!9w4w!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d553155-7dc6-4d4d-9fee-ce18f21006e9_1600x405.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Figure: &nbsp;arxiv.1904.09675</em></p><p>To calculate the BERTScore in Python, you can use the bert_score library. You can install it using:</p><pre><code>pip install torch torchvision torchaudio

pip install bert-score

Here&#8217;s an example of how to use the library to calculate BERTScore:

import torch

from bert_score import score

cands = [''The quick brown dog jumps over the lazy fox.'']

refs = [''The quick brown fox jumps over the lazy dog.'']

P, R, F1 = score(cands, refs, lang=''en'', verbose=True)

print(F1)</code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VhSt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VhSt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png 424w, https://substackcdn.com/image/fetch/$s_!VhSt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png 848w, https://substackcdn.com/image/fetch/$s_!VhSt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png 1272w, https://substackcdn.com/image/fetch/$s_!VhSt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VhSt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png" width="1072" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1072,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:31778,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VhSt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png 424w, https://substackcdn.com/image/fetch/$s_!VhSt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png 848w, https://substackcdn.com/image/fetch/$s_!VhSt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png 1272w, https://substackcdn.com/image/fetch/$s_!VhSt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dd063f-1589-49e4-9994-b2deee334a2d_1072x344.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Langchain Evaluation</h3><p>LangChain offers various types of evaluators to help you measure performance and integrity on diverse data, and we hope to encourage the community to create and share other useful evaluators so everyone can improve. These docs will introduce the evaluator types, how to use them, and provide some examples of their use in real-world scenarios. These built-in evaluators all integrate smoothly with <a href="https://python.langchain.com/v0.1/docs/langsmith/">LangSmith</a>, and allow you to create feedback loops that improve your application over time and prevent regressions.</p><p>Each evaluator type in LangChain comes with ready-to-use implementations and an extensible API that allows for customization according to your unique requirements. Here are some of the types of evaluators we offer:</p><ul><li><p><strong><a href="https://python.langchain.com/v0.1/docs/guides/productionization/evaluation/string/">String Evaluators</a>:</strong> These evaluators assess the predicted string for a given input, usually comparing it against a reference string.</p></li><li><p><strong><a href="https://python.langchain.com/v0.1/docs/guides/productionization/evaluation/trajectory/">Trajectory Evaluators</a>:</strong> These are used to evaluate the entire trajectory of agent actions.</p></li><li><p><strong><a href="https://python.langchain.com/v0.1/docs/guides/productionization/evaluation/comparison/">Comparison Evaluators</a>: </strong>These evaluators are designed to compare predictions from two runs on a common input.</p></li></ul><p>Learn more: https://python.langchain.com/v0.1/docs/guides/productionization/evaluation/</p><h1>Ethical considerations in LLM development</h1><p>Ethical considerations in the development and deployment of Large Language Models (LLMs) like GPT-4 are crucial due to their potential societal impacts. Here are eight key ethical concerns:</p><p><strong>1. Generating Harmful Content: </strong>LLMs can inadvertently generate content that promotes hate speech, extremism, or discrimination, reflecting biases present in their training data. This can lead to societal issues such as incitement to violence or social unrest.</p><p><strong>2. Economic Impact: </strong>As LLMs become more widespread and powerful, they can disrupt the job market by automating certain tasks, leading to workforce displacement and exacerbating inequality. It's essential to develop policies that promote technical literacy and address these impacts.</p><p><strong>3. Hallucinations: </strong>LLMs may produce false or misleading information, which can be problematic as they become more convincing. It's crucial to train these models on accurate and relevant datasets to minimize hallucinations.</p><p><strong>4. Disinformation &amp; Influencing Operations: </strong>LLMs can spread disinformation or be used by bad actors for influence operations, potentially impacting public opinion and policy. Developing fact-checking mechanisms and enhancing media literacy are necessary to counter this.</p><p><strong>5. Weapon Development: </strong>There's a concern that LLMs could be used to gather information about weapons production, posing a security risk. Implementing security measures is essential to prevent misuse.</p><p><strong>6. Privacy: </strong>LLMs require access to large amounts of data, including personal information, which can lead to privacy concerns. Clear policies on data collection and storage, and the practice of data anonymization, are necessary to address these issues.</p><p><strong>7. Risky Emergent Behaviors: </strong>LLMs may exhibit unpredictable behaviors, such as formulating long-term plans or striving for authority, which can be risky, especially when interacting with other systems. Measures should be put in place to mitigate these risks.</p><p><strong>8. Unwanted Acceleration: </strong>LLMs can accelerate innovation and scientific discovery, potentially leading to a race in AI development that might undermine safety and ethical standards. There's a call for a moratorium on developing more powerful AI systems to address this concern.</p><p>Addressing these ethical considerations requires a multifaceted approach, including developing and deploying LLMs responsibly, promoting technical literacy, and implementing robust security measures and ethical guidelines.</p><h1>Debugging techniques for LLMs</h1><p>Debugging techniques for Large Language Models (LLMs) can be categorized into three main areas: visualizing attention, gradient analysis, and ablation studies. These techniques help in understanding the model's behavior, identifying issues, and optimizing the model's performance.</p><h2>Visualizing Attention</h2><p>Visualizing attention is crucial for understanding how LLMs process and generate text. Tools like BertViz provide interactive visualizations of attention mechanisms in models like BERT, GPT-2, or T5. These visualizations can help identify which parts of the input text the model is focusing on, and how it is computing attention across different layers and heads.</p><p>For example, using BertViz, you can generate HTML representations of the attention mechanism for specific inputs. This can be done by setting the `html_action` parameter to `'return'` in the `head_view` or `neuron_view` functions, which allows you to save the visualization as an HTML file or process it further in a Python environment.</p><pre><code>from transformers import AutoTokenizer, AutoModel, utils

from bertviz import head_view

utils.logging.set_verbosity_error()&nbsp; # Suppress standard warnings

tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

model = AutoModel.from_pretrained("bert-base-uncased", output_attentions=True)

inputs = tokenizer.encode("the rabbit quickly hopped the turtle slowly crawled", return_tensors='pt')

outputs = model(inputs)

attention = outputs[-1]&nbsp; # Output includes attention weights when output_attentions=True

tokens = tokenizer.convert_ids_to_tokens(inputs[0])&nbsp;&nbsp;

html_head_view = head_view(attention, tokens, html_action='return')

with open("PATH_TO_YOUR_FILE/head_view.html", 'w') as file:

&nbsp;&nbsp;&nbsp;&nbsp;file.write(html_head_view.data)</code></pre><p>Result as follow:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GuSF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GuSF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png 424w, https://substackcdn.com/image/fetch/$s_!GuSF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png 848w, https://substackcdn.com/image/fetch/$s_!GuSF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png 1272w, https://substackcdn.com/image/fetch/$s_!GuSF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GuSF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png" width="1174" height="1022" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1022,&quot;width&quot;:1174,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GuSF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png 424w, https://substackcdn.com/image/fetch/$s_!GuSF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png 848w, https://substackcdn.com/image/fetch/$s_!GuSF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png 1272w, https://substackcdn.com/image/fetch/$s_!GuSF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc22c539c-81df-4037-9aa5-44d7a7bb5832_1174x1022.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Do install bertviz and tokenize sample sentences.</p><pre><code>!pip install bertviz

# Load model and retrieve attention weights

from bertviz import head_view, model_view

from transformers import BertTokenizer, BertModel

model_version = 'bert-base-uncased'

model = BertModel.from_pretrained(model_version, output_attentions=True)

tokenizer = BertTokenizer.from_pretrained(model_version)

sentence_a = "The cat sat on the mat"

sentence_b = "The cat lay on the rug"

inputs = tokenizer.encode_plus(sentence_a, sentence_b, return_tensors='pt')

input_ids = inputs['input_ids']

token_type_ids = inputs['token_type_ids']

attention = model(input_ids, token_type_ids=token_type_ids)[-1]

sentence_b_start = token_type_ids[0].tolist().index(1)

input_id_list = input_ids[0].tolist() # Batch index 0

tokens = tokenizer.convert_ids_to_tokens(input_id_list)&nbsp;</code></pre><h3>Head View</h3><p>The head view visualizes attention in one or more heads from a single Transformer layer. Each line shows the attention from one token (left) to another (right). Line weight reflects the attention value (ranges from 0 to 1), while line color identifies the attention head.When multiple heads are selected (indicated by the colored tiles at the top), the corresponding visualizations are overlaid onto one another.&nbsp;</p><p>head_view(attention, tokens, sentence_b_start)</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MggY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MggY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png 424w, https://substackcdn.com/image/fetch/$s_!MggY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png 848w, https://substackcdn.com/image/fetch/$s_!MggY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png 1272w, https://substackcdn.com/image/fetch/$s_!MggY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MggY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png" width="850" height="986" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:986,&quot;width&quot;:850,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MggY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png 424w, https://substackcdn.com/image/fetch/$s_!MggY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png 848w, https://substackcdn.com/image/fetch/$s_!MggY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png 1272w, https://substackcdn.com/image/fetch/$s_!MggY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0331e58e-7224-456e-b1ce-7b9246682a95_850x986.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Model View</h3><p>The model view provides a birds-eye view of attention throughout the entire model. Each cell shows the attention weights for a particular head, indexed by layer (row) and head (column). The lines in each cell represent the attention from one token (left) to another (right), with line weight proportional to the attention value (ranges from 0 to 1). Use below line to visualize it:</p><pre><code>model_view(attention, tokens, sentence_b_start)</code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D_Pv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D_Pv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png 424w, https://substackcdn.com/image/fetch/$s_!D_Pv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png 848w, https://substackcdn.com/image/fetch/$s_!D_Pv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png 1272w, https://substackcdn.com/image/fetch/$s_!D_Pv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D_Pv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png" width="1448" height="1246" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1246,&quot;width&quot;:1448,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D_Pv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png 424w, https://substackcdn.com/image/fetch/$s_!D_Pv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png 848w, https://substackcdn.com/image/fetch/$s_!D_Pv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png 1272w, https://substackcdn.com/image/fetch/$s_!D_Pv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fd6649a-4e67-4b9e-bd18-b1a90e3ccdaa_1448x1246.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This technique is particularly useful for understanding how attention is distributed across different parts of the input text and how it influences the model's output.</p><h3>Neuron View</h3><p>The neuron view visualizes the intermediate representations (e.g. query and key vectors) that are used to compute attention. In the collapsed view (initial state), the lines show the attention from each token (left) to every other token (right).Here is the example:</p><pre><code>from bertviz.transformers_neuron_view import BertModel, BertTokenizer

from bertviz.neuron_view import show

model_type = 'bert'

model_version = 'bert-base-uncased'

model = BertModel.from_pretrained(model_version, output_attentions=True)

tokenizer = BertTokenizer.from_pretrained(model_version, do_lower_case=True)

show(model, model_type, tokenizer, sentence_a, sentence_b, layer=4, head=3)</code></pre><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MCbA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MCbA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png 424w, https://substackcdn.com/image/fetch/$s_!MCbA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png 848w, https://substackcdn.com/image/fetch/$s_!MCbA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png 1272w, https://substackcdn.com/image/fetch/$s_!MCbA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MCbA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png" width="1034" height="1118" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1118,&quot;width&quot;:1034,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MCbA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png 424w, https://substackcdn.com/image/fetch/$s_!MCbA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png 848w, https://substackcdn.com/image/fetch/$s_!MCbA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png 1272w, https://substackcdn.com/image/fetch/$s_!MCbA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79db2d1a-173e-4da5-875e-f8b4d9febe84_1034x1118.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Gradient Analysis</h2><p>Gradient analysis involves examining the gradients of the model's parameters during training or inference. This can help identify issues such as vanishing or exploding gradients, which can affect the model's learning process and performance. By visualizing gradients, developers can gain insights into which parts of the model are learning effectively and which may require adjustments.</p><h3>GradSafe</h3><p>Gradient Analysis for Large Language Models (LLMs) involves a novel approach to detecting unsafe prompts by analyzing the gradients of safety-critical parameters in LLMs. This method, known as GradSafe, is grounded in the observation that the gradients of an LLM&#8217;s loss for unsafe prompts paired with a compliance response exhibit similar patterns on certain safety-critical parameters, in contrast to the divergent patterns observed with safe prompts. These patterns are used to accurately detect unsafe prompts without the need for extensive data collection or training processes.The process involves several key steps:</p><p><strong>Identifying Safety-Critical Parameters: </strong></p><p>The first step involves identifying safety-critical parameters where gradients derived from unsafe prompts and safe prompts can be distinguished. This is based on the conjecture that the gradients of an LLM&#8217;s loss for pairs of unsafe prompt and compliance response (such as &#8216;Sure&#8217;) on the safety-critical parameters are expected to manifest similar patterns. Conversely, similar effects are not anticipated for a pair of safe prompt and compliance response.</p><p><strong>Computing Gradients and Cosine Similarities: </strong></p><p>For each gradient matrix, slices are made both row-wise and column-wise to identify safety-critical parameters and calculate cosine similarity features. The average of the gradient slices for all unsafe prompts serves as reference gradient slices for subsequent cosine similarity computations. The aim is to identify parameter slices exhibiting high similarity in gradients across unsafe prompts, while demonstrating low similarity between unsafe and safe prompts.</p><p><strong>Detecting Unsafe Prompts: </strong></p><p>GradSafe evaluates the safety of a prompt by comparing its gradients of safety-critical parameters, when paired with a compliance response, with the unsafe gradient reference. Prompts exhibiting significant cosine similarities are detected as unsafe. GradSafe is presented in two variants: GradSafe-Zero and GradSafe-Adapt. GradSafe-Zero relies solely on the cosine similarity averaged across all safety-critical parameters to determine whether a prompt is unsafe. GradSafe-Adapt, on the other hand, undergoes adjustments by training a simple logistic regression model with cosine similarities as features, leveraging the training set to facilitate domain adaptation.</p><p><strong>Performance Evaluation: </strong></p><p>The performance of GradSafe is evaluated using datasets like ToxicChat and XSTest, with metrics such as Area Under the Precision-Recall Curve (AUPRC), precision, recall, and F1 scores. The results demonstrate that GradSafe, applied to Llama-2 without further training, outperforms Llama Guard, despite its extensive finetuning with a large dataset, in detecting unsafe prompts. This superior performance is consistent across both zero-shot and adaptation scenarios.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s-vt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s-vt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png 424w, https://substackcdn.com/image/fetch/$s_!s-vt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png 848w, https://substackcdn.com/image/fetch/$s_!s-vt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png 1272w, https://substackcdn.com/image/fetch/$s_!s-vt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s-vt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png" width="768" height="1158" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1158,&quot;width&quot;:768,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!s-vt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png 424w, https://substackcdn.com/image/fetch/$s_!s-vt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png 848w, https://substackcdn.com/image/fetch/$s_!s-vt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png 1272w, https://substackcdn.com/image/fetch/$s_!s-vt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8bbe72d-bba0-4f2d-96dc-17900bb828f7_768x1158.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Ablation Studies</h2><p>Ablation studies involve systematically removing or modifying components of the model to understand their impact on performance. This can include removing entire layers, changing the model's architecture, or altering the training data. By comparing the model's performance before and after these changes, developers can identify which components are most critical to the model's performance and where potential improvements can be made.</p><h1>Latest research trends in fine-tuning, adaptation, evaluation, and debugging of LLMs</h1><p>The latest research trends in fine-tuning, adaptation, evaluation, and debugging of Large Language Models (LLMs) encompass a wide range of innovative approaches and methodologies to enhance their performance, reliability, and applicability across various domains. Here are some key trends and findings:</p><h4>Fine-Tuning and Adaptation</h4><p><strong>- Domain-Specific Fine-Tuning: </strong>Research emphasizes the importance of fine-tuning LLMs on domain-specific data to improve their performance on tasks related to that domain. This includes exposing the model to specialized datasets that can enhance its accuracy, relevance, and effectiveness for intended use cases.</p><p><strong>- Instruction-Fine-Tuned Models:</strong> A notable trend is the scaling of instruction-fine-tuned language models, which adapt LLMs to perform tasks based on explicit instructions. This approach allows for more controlled and targeted training, enhancing the model's ability to perform specific tasks with high accuracy.</p><h4>Evaluation</h4><p><strong>- Automatic Evaluation Frameworks:</strong> The development of frameworks like FLASK (Fine-Grained Language Model Evaluation Based on Alignment Skill Sets) and INSTRUCTSCORE aims to provide explainable and reliable evaluation metrics for LLMs. These frameworks focus on evaluating LLMs based on their alignment with human judgments, offering insights into the models' strengths and weaknesses.</p><p><strong>- Multi-Agent Debate Evaluation: </strong>Chateval introduces a method for evaluating LLMs through multi-agent debate, which simulates a conversation between multiple models to assess their performance in generating coherent and contextually relevant responses.</p><p><strong>- Automatic Dialogue Evaluation: </strong>Studies explore the use of LLMs as automatic dialogue evaluators, assessing their ability to understand and respond to dialogues effectively. This includes evaluating malevolence in dialogues and using LLMs to judge the quality of generated text.</p><h3>Debugging</h3><p><strong>- Interpreting Models with Contrastive Explanations: </strong>Research on interpreting language models with contrastive explanations focuses on understanding how LLMs generate text and the factors that influence their output. This includes identifying patterns and biases in the models' outputs.</p><p><strong>- Cross-Examination for Factual Error Detection:</strong> The concept of using LLMs for cross-examination to detect factual errors in text has been explored, highlighting the potential of LLMs in verifying the accuracy of information.</p><h3>General Trends</h3><p><strong>- Scalability and Fairness: </strong>There's a trend towards developing wider and deeper LLM networks that can serve as fairer evaluators. This includes exploring the scalability of fine-tuned LLMs as judges, indicating a move towards more efficient and scalable solutions.</p><p><strong>- Comprehensive Analysis of LLM Effectiveness: </strong>A comprehensive analysis of the effectiveness of LLMs as automatic dialogue evaluators and the development of tools like MT-Bench and Chatbot Arena for judging LLMs-as-judges have been conducted, aiming to assess their capabilities in various evaluation scenarios.</p><h1>Summary</h1><p>This chapter discusses techniques for adapting, evaluating, and debugging Large Language Models (LLMs) to enhance their performance on specific tasks. It covers fine-tuning methods such as task-specific heads, unfreezing subsets of parameters (PEFT), and various advanced fine-tuning techniques including RLHF-based, DPO, CPL, RLAIF, HITL, Prompt-based Learning, LfD, and Meta-learning. The chapter also explores domain adaptation and transfer learning through intermediate pre-training and data augmentation. For continuous learning and model updates, it outlines the importance of these practices. Evaluation metrics like Perplexity, accuracy, F1, BLEU, and human evaluations are discussed, followed by ethical considerations in LLM development. Debugging techniques include visualizing attention, gradient analysis, and ablation studies. Finally, it touches on the latest research trends in fine-tuning, adaptation, evaluation, and debugging of LLMs.</p><h2>Quiz questions</h2><p>Here is a quiz to assess understanding of this chapter :</p><p>1. Which of the following is a technique for fine-tuning LLMs on specific tasks?</p><p>A) Unfreezing subsets of parameters - PEFT</p><p>B) RLHF-based fine-tuning</p><p>C) Learning from Demonstrations (LfD)</p><p>D) All of the above</p><p>2. Which technique involves modifying a pre-trained model to perform better on a new task by adjusting its parameters?</p><p>A) Task-specific heads</p><p>B)&nbsp; Intermediate pre-training</p><p>C) &nbsp; Direct Preference Optimization (DPO)</p><p>D) &nbsp; Data augmentation</p><p>3. What is a common evaluation metric used to assess the performance of LLMs on downstream tasks?</p><p>A) Perplexity</p><p>B) F1 score</p><p>C) BLEU score</p><p>D) All of the above</p><p>4. Which technique is used to debug LLMs by analyzing the model's learning process?</p><p>A) Visualizing attention</p><p>B) Gradient analysis</p><p>C)&nbsp; Ablation studies</p><p>D)&nbsp; All of the above</p><p>5. In the context of LLM performance testing, what does the Langchain evaluator return if the LLM meets the defined criteria?</p><p>A) 1</p><p>B) 0</p><p>C)&nbsp; -1</p><p>D)&nbsp; None of the above</p><p>6. What is the primary goal of creating a 'gold test set' in LLM evaluation?</p><p>A) To measure the LLM's performance against actual production data</p><p>B) To test the LLM against a wide range of tasks and scenarios</p><p>C)&nbsp; To provide a benchmark against which all responses are compared</p><p>D)&nbsp; To ensure the LLM is tested against realistic challenges</p><p>7. Which library offers a suite of scoring metrics for evaluating and comparing LLMs, prompts, and hyperparameters, aiming to assist in making data-driven decisions?</p><p>A) UpTrain</p><p>B) Deep-Eval</p><p>C) Arthur Bench</p><p>D) RAGAS</p><p></p><p><strong>Correct Answers:</strong></p><p>1.&nbsp; D. All of the above</p><p>2.&nbsp; A. Task-specific heads</p><p>3.&nbsp; D. All of the above</p><p>4.&nbsp; D. All of the above</p><p>5.&nbsp; A. 1</p><p>6.&nbsp; C. To provide a benchmark against which all responses are compared</p><p>7.&nbsp; C. Arthur Bench</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Chapter 4 - Introduction to Large Language Models ]]></title><description><![CDATA[Overview of LLMs]]></description><link>https://aboniasojasingarayar.substack.com/p/introduction-to-large-language-models</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/introduction-to-large-language-models</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Mon, 26 Aug 2024 07:15:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Mbz5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As you flip through the pages of this chapter, get ready to delve into the heart of these game-changing AI marvels. We'll embark on a journey to understand what LLMs are, how they work, and the vast potential they hold for transforming various aspects of our lives. But before we dive in, let's set the stage with a simple question: what exactly are LLMs? Imagine a language model trained on a massive dataset of text and code, capable of generating human-quality text, translating languages, writing different kinds of creative content, and even answering your questions in an informative way. That's the essence of an LLM &#8211; that can process and understand information like never before.</p><p>This chapter will be your comprehensive guide to navigating the fascinating world of LLMs. We'll delve into their core concepts, exploring different types like autoregressive models and encoder-decoder models. You'll discover the magic behind self-attention, a mechanism that allows LLMs to focus on relevant information, and delve into the pre-training strategies that give them their vast knowledge. Finally, we'll showcase the real-world applications of LLMs, from powering chatbots and generating realistic dialogue to creating marketing copy and summarizing complex topics.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>By the end of this chapter, you'll have a comprehensive understanding of LLMs, their inner workings, and the remarkable ways they are shaping the future of language and communication. So, prepare to be amazed by the power of these linguistic powerhouses!</p><p>In this chapter, we will cover the following topics :</p><ul><li><p>Definition and overview of LLMs</p></li><li><p>Types of LLMs</p></li><li><p>Key technical concepts</p></li><li><p>Evaluating LLMs</p></li><li><p>Applications, Challenges Limitation</p></li><li><p>LLM Cheat Sheet</p></li></ul><h2>Definition and overview of LLMs</h2><p>Imagine machines that converse like seasoned storytellers, translate languages with seamless fluency, and even craft poems that tug at your heartstrings. These aren't characters from science fiction, but a reality ushered in by a new breed of artificial intelligence: Large Language Models (LLMs).</p><p>LLMs are more than just advanced chatbots; they're sophisticated AI models trained on colossal datasets of text and code. Think of them as language titans, trained in libraries filled with novels, news articles, and even code repositories. This vast exposure allows them to grasp the nuances of human language, understand context, and even generate their own creative text formats.</p><p>But how do these titans function? At their core lies a complex neural network architecture called a transformer. Imagine this as a network of interconnected nodes, each processing and analyzing bits of information. Like a team of detectives piecing together clues, the network deciphers the relationships between words, learns their meanings, and ultimately predicts the next word in a sequence. This allows LLMs to understand the flow of language, generate responses that make sense, and even perform impressive feats like translating languages or summarizing text.</p><p>LLMs are no longer confined to research labs. They're actively shaping various industries, with applications ranging from generating personalized learning experiences in education to powering chatbots that offer natural language interactions in customer service. Imagine chatbots that understand your frustration and respond with genuine empathy, or AI assistants that summarize research papers in clear and concise language. These are just glimpses of the potential LLMs hold.</p><p>However, like any powerful tool, LLMs come with their own set of challenges. We need to be mindful of potential biases, the spread of misinformation, and the ethical considerations surrounding this technology. Responsible development and deployment are crucial to ensure LLMs are a force for good, amplifying human creativity and understanding while navigating the ethical landscape with care.As we delve deeper into the world of LLMs, remember this: they're not merely machines mimicking human language; they represent a fascinating intersection of artificial intelligence and human creativity, with the potential to reshape the way we interact with information and express ourselves. This journey has just begun, and the future holds exciting possibilities for the language titans we call LLMs.</p><p>Below chart shows the timeline of some of the most representative LLM frameworks (so far). In addition to large language models with our #parameters threshold, we included a few representative works, which pushed the limits of language models, and paved the way for their success (e.g. vanilla Transformer, BERT, GPT-1), as well as some small language models. &#9827; shows entities that serve not only as models but also as approaches. &#9830; shows only approaches.[LLM Survey arxiv.2402.06196]</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Mbz5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Mbz5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png 424w, https://substackcdn.com/image/fetch/$s_!Mbz5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png 848w, https://substackcdn.com/image/fetch/$s_!Mbz5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png 1272w, https://substackcdn.com/image/fetch/$s_!Mbz5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Mbz5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png" width="1456" height="691" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:691,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Mbz5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png 424w, https://substackcdn.com/image/fetch/$s_!Mbz5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png 848w, https://substackcdn.com/image/fetch/$s_!Mbz5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png 1272w, https://substackcdn.com/image/fetch/$s_!Mbz5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c943d1d-60ff-47b4-a562-4c8288ef3f84_1600x759.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Notable LLMs</h3><p>The journey of LLMs hasn't been a straight shot. The story starts with simpler statistical models, followed by the groundbreaking Generative Pre-trained Transformer (GPT) in 2017. This was a turning point, introducing the transformer architecture that empowered more powerful LLMs. Then came 2018 with BERT, setting new records in NLP tasks and showcasing the power of pre-training, where models are exposed to vast amounts of text before tackling specific tasks. Finally, 2020 witnessed the arrival of GPT-3, a massive autoregressive LLM capable of generating remarkably human-like text, capturing the public's imagination and highlighting the potential and challenges of this technology.</p><p>The landscape of LLMs has rapidly evolved since 2020. These models range from web interfaces and APIs to completely accessible datasets, codebases, and model checkpoints. It is important to note that the terms for their commercial usage also differ. In this section, we highlight notable LLM models in chronological order, showcasing their unique features and contributions.</p><p>GPT-3 [API] was released by OpenAI in June 2020. The model contains 175 billion parameters and is considered one of the most important LLM milestones. It was the first model to demonstrate strong few-shot learning capabilities.</p><p>ChatGPT/GPT-3.5 [API] was released by OpenAI in November 2022. It extended GPT-3 in terms of both the data and underlying algorithms. For example, it used RLHF to boost model performance. ChatGPT was specifically designed for conversational use, offering a more natural and engaging chatbot experience.</p><p>GPT-4 [API] was released in March 2023. In addition to handling text, it can also take images as input. GPT-4 outperformed ChatGPT on many tasks, including passing the Bar exam. It also demonstrated better performance at safety tests (OpenAI, 2023).</p><p>LLaMA [checkpoint] was released by Meta in Feb 2023. It is a group of models with different numbers of parameters (7B, 13B, 33B, and 65B). Unlike GPT-X models that are accessible only through APIs, LLaMA provides implementation details of the model architecture and checkpoints. It has become a fundamental resource for subsequent open-source research..</p><p>Alpaca [checkpoint] was released by Stanford in March 2023. Alpaca was fine-tuned on top of LLaMA-7B using data generated by GPT-3. Remarkably, Alpaca demonstrated comparable performance to much larger models like GPT-3 (175B) while being smaller and more cost-effective to reproduce (less than $600).</p><p>Claude[API] was released by Alphabet-backed company Anthropic in March 2023. By adopting the Constitutional AI method proposed by Yuntao et al. (2022), Claude is reported to generate less toxic, biased, and hallucinatory responses. In May 2023, Anthropic expanded Claude&#8217;s context window from 9k to 100K. This is equivalent to around 75,000 words and allows the model to parse all the information from an entire book at once.</p><p>Falcon[checkpoint] was released by TII in May 2023. It is an open-source LLM with 40 billion parameters under the Apache 2.0 license. Falcon stands out due to the high quality of its training data, which includes 1,000 billion tokens from the RefinedWeb enhanced with curated corpora. At the time of writing, Falcon holds the top position on the Huggingface Open LLM Leaderboard.</p><p>Gemini Ultra released in 2024 , the version with Ultra will be called Gemini Advanced with a 540-billion parameter model, a new experience far more capable at reasoning, following instructions, coding, and creative collaboration. For example, it can be a personal tutor, tailored to your learning style. Or it can be a creative partner, helping you plan a content strategy or build a business plan.</p><h3>NLP and its Connection to LLMs</h3><p>Before we fully appreciate the magic of LLMs, we need to understand the language game they're playing: Natural Language Processing (NLP). Imagine NLP as the bridge between computers and the messy, wonderful world of human language. It's the art of teaching machines to understand, interpret, and even generate human language &#8211; a complex task filled with nuances and challenges.</p><p>At the heart of NLP lie various tasks, each posing its own unique hurdles. Let's explore some of the core areas:</p><p>Text classification: Can a machine categorize a document as news, a poem, or an email? LLMs tackle this by analyzing the text style and content.</p><p>Machine translation: Transforming words from one language to another accurately and preserving meaning is no easy feat. LLMs bridge this gap, understanding the context and intent to deliver fluent translations.</p><p>Question answering: Imagine asking a machine a question and getting a clear, informative answer. LLMs dive into vast amounts of information, seeking the answer you need amidst the sea of text.</p><p>Text summarization: Condensing a lengthy article into a concise summary without losing key points? LLMs are adept at identifying the main ideas and presenting them in a clear, digestible format.</p><p>Text generation: From crafting poems to writing scripts, LLMs can even generate creative text formats, pushing the boundaries of machine-made language.</p><p>These are just a few examples, and the list continues to grow. But what makes LLMs particularly adept at tackling these NLP challenges? Their secret lies in their training. Like master linguists, they're exposed to massive amounts of text and code, learning the intricacies of language through this exposure. This allows them to identify patterns, understand context, and ultimately perform these NLP tasks with impressive accuracy and creativity.</p><p>Think of it this way: imagine training a human translator by immersing them in different cultures and languages for years. That's essentially what happens with LLMs, except on a vastly accelerated scale. This training empowers them to tackle the complexities of NLP, pushing the boundaries of what machines can achieve with language.</p><p>In the following section, we'll delve deeper into how LLMs are used for specific NLP applications, showcasing their capabilities and exploring the exciting possibilities they hold for the future. Remember, the connection between NLP and LLMs is vital &#8211; understanding the tasks and challenges helps us appreciate the true power and potential of these language titans.</p><h2>Exploring Different Types of LLMs</h2><p>Now that we've peeked into the world of NLP and how LLMs use it, let's meet the different types of these language magicians! we explore two main categories:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1re_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1re_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png 424w, https://substackcdn.com/image/fetch/$s_!1re_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png 848w, https://substackcdn.com/image/fetch/$s_!1re_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png 1272w, https://substackcdn.com/image/fetch/$s_!1re_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1re_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png" width="1418" height="958" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:958,&quot;width&quot;:1418,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:98405,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1re_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png 424w, https://substackcdn.com/image/fetch/$s_!1re_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png 848w, https://substackcdn.com/image/fetch/$s_!1re_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png 1272w, https://substackcdn.com/image/fetch/$s_!1re_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57e8287a-c95a-4d94-a3e1-7e3c41472789_1418x958.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Autoregressive Models</h3><p>Autoregressive Language Models (e.g., GPT): Autoregressive models primarily use the decoder part of the Transformer architecture, making them well-suited for natural language generation (NLG) tasks like text summarization, generation, etc. These models generate text by predicting the next word in a sequence given the previous words. They are trained to maximize the likelihood of each word in the training dataset, given its context. The most well-known example of an autoregressive language model is OpenAI&#8217;s GPT (Generative Pre-trained Transformer) series, with GPT-4 being the latest and most powerful iteration. Autoregressive models based on decoder networks primarily leverage layers related to self-attention, cross-attention mechanisms, and feed-forward networks as part of their neural network architecture. It works by analyzing the words they've already created and predict what word would likely come next, based on their training data. Think of it as guessing the next word in a game of Mad Libs, but with way more complex calculations! Their training involves massive amounts of text, like books, articles, and even code. This exposure helps them understand language patterns and predict words that make sense, leading to surprisingly human-like text generation.</p><p>Specifically, given a text sequence<em>&nbsp;</em></p><div class="pullquote"><p><em>x=(x1,&#8943;,xT)</em></p></div><p>AR language modeling factorizes the likelihood into a forward product<em>&nbsp;</em></p><div class="pullquote"><p><em>p(x)=&#8719;Tt=1p(xt&#8739;x&lt;t)</em></p></div><p>or a backward one&nbsp;</p><div class="pullquote"><p><em>p(x)=&#8719;1t=Tp(xt&#8739;x&gt;t)</em></p></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cu-B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cu-B!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png 424w, https://substackcdn.com/image/fetch/$s_!cu-B!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png 848w, https://substackcdn.com/image/fetch/$s_!cu-B!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png 1272w, https://substackcdn.com/image/fetch/$s_!cu-B!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cu-B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png" width="1404" height="798" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:798,&quot;width&quot;:1404,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:57080,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!cu-B!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png 424w, https://substackcdn.com/image/fetch/$s_!cu-B!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png 848w, https://substackcdn.com/image/fetch/$s_!cu-B!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png 1272w, https://substackcdn.com/image/fetch/$s_!cu-B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd97ea933-cd50-4952-a566-57d135f41cb9_1404x798.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Strengths and Limitations</h4><p>They excel at storytelling, poem writing, and generating different creative text formats. Imagine them as AI bards, composing lyrics or scripts with impressive fluency. They can adjust their "voice" and style based on the context and prompts they receive. Think of them as chameleons of language, adapting to different genres and tones. AR language models are adept at generative NLP tasks because they utilize causal attention to predict the next token, making them naturally suited for content generation. One advantage of AR models is that generating data for them is relatively straightforward, as the training objective can simply be to predict the next token in a given corpus. However, it's important to note that while AR models consider either forward or backward context for each token prediction, they still incorporate bidirectional context by conditioning predictions on the entire sequence generated up to that point.</p><h4>Popular autoregressive models</h4><p>GPT, GPT-2, GPT-3, and CTRL: These models&nbsp; gained fame for their remarkable text generation capabilities.</p><p>Jukebox: Imagine creating music just by describing it! Jukebox uses the power of autoregression to compose melodies and lyrics.</p><p>Megatron-Turing NLG: This large model focuses on generating different creative text formats, like poems and code.</p><h3>Autoencoding Language Models&nbsp;</h3><p>Autoencoding Language Models (e.g., BERT): Autoencoding models, on the other hand, mainly use the encoder part of the Transformer. It&#8217;s designed for tasks like classification, question answering, etc. These models learn to generate a fixed-size vector representation (also called embeddings) of input text by reconstructing the original input from a masked or corrupted version of it. They are trained to predict missing or masked words in the input text by leveraging the surrounding context. BERT (Bidirectional Encoder Representations from Transformers), developed by Google, is one of the most famous autoencoding language models. It can be fine-tuned for a variety of NLP tasks, such as sentiment analysis, named entity recognition and question answering. Autoencoding models based on encoder networks primarily leverage layers related to self-attention mechanisms and feed-forward networks as part of their neural network architecture.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bA4D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bA4D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png 424w, https://substackcdn.com/image/fetch/$s_!bA4D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png 848w, https://substackcdn.com/image/fetch/$s_!bA4D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png 1272w, https://substackcdn.com/image/fetch/$s_!bA4D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bA4D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png" width="1438" height="408" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:408,&quot;width&quot;:1438,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38206,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bA4D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png 424w, https://substackcdn.com/image/fetch/$s_!bA4D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png 848w, https://substackcdn.com/image/fetch/$s_!bA4D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png 1272w, https://substackcdn.com/image/fetch/$s_!bA4D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2f99e44-9c20-4909-969c-44b2c731c183_1438x408.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>&nbsp;</p><p>AELMs take a sentence as input and compress it into a smaller, latent representation. Think of it as summarizing the essence of the sentence in a condensed code. This code captures the key information and relationships between words.Once they have the compressed code, AELMs possess the remarkable ability to decode it back into a new sentence. This reconstructed sentence may not be identical to the original, but it retains the core meaning and conveys the same message.A notable example is BERT, which has been the state-of-the-art pretraining approach. Given the input token sequence, a certain portion of tokens are replaced by a special symbol [MASK], and the model is trained to recover the original tokens from the corrupted version. The AE language model aims to reconstruct the original data from corrupted input.</p><h4>Strengths&nbsp;</h4><p>1. Unsupervised Learning: AE models excel in unsupervised learning scenarios, where they autonomously learn the intricate patterns and structures of language from vast amounts of unannotated text data. This capability renders them remarkably versatile, capable of adaptation to a myriad of tasks without the constraints of task-specific labeled datasets.</p><p>2. Contextual Understanding: By reconstructing input sequences from corrupted or masked tokens, AE models cultivate a deep understanding of contextual relationships within language. This contextual awareness enables them to capture nuanced meanings and dependencies between words or tokens, facilitating more accurate and contextually relevant outputs.</p><p>3. Pretraining-Fine Tuning: AE models operate on a pretraining-finetuning paradigm, wherein they undergo initial training on large corpora of text data followed by fine-tuning on task-specific datasets. This approach facilitates transfer learning, where knowledge gleaned during pretraining is seamlessly transferred to downstream tasks, resulting in enhanced performance, particularly in scenarios with limited task-specific data.</p><p>4. Flexible Architecture: AE models offer flexibility in architecture design, ranging from Transformer-based structures like BERT to autoencoder architectures such as ALBERT. This adaptability empowers researchers to experiment with diverse architectures, tailoring them to specific use cases or computational resources.</p><h4>Limitations</h4><p>1. Masking-based Pretraining: AE models rely on token masking during pre training, which can introduce a discrepancy between the pre training and fine tuning phases. As masked tokens are absent during finetuning, this misalignment may impact performance on downstream tasks, necessitating careful consideration during model development.</p><p>2. Limited Context Window: AE models typically operate within a fixed context window size, constraining their ability to capture long-range dependencies in language. This limitation may lead to suboptimal performance, particularly in tasks requiring comprehension of extended contexts.</p><p>3. Fine-tuning Data Requirement: While AE models leverage transfer learning to enhance performance, achieving optimal results often necessitates a substantial amount of task-specific data for fine-tuning. In scenarios with limited or low-quality labeled data, the full benefits of pretraining may not be realized, requiring strategic resource allocation and data augmentation strategies.</p><p>4. Semantic Understanding: While proficient in capturing syntactic patterns and contextual relationships, AE models may encounter challenges in deeper semantic understanding, such as reasoning or inference. This limitation may constrain their performance on tasks demanding higher-level language comprehension and reasoning abilities.</p><h4>Notable AELMs&nbsp;</h4><p>BERT: This pioneer paved the way for AELMs by demonstrating their effectiveness in various NLP tasks.</p><p>BART: This versatile model excels at text summarization, translation, and even question answering.</p><p>XLNet: This powerful model uses a unique permutation language modeling approach, pushing the boundaries of AELM capabilities.</p><h3>Encoder-decoder models</h3><p>The third one is the combination of autoencoding and autoregressive such as the T5 (Text-to-Text Transfer Transformer) model. Developed by Google in 2020, T5 LLM can perform natural language understanding (NLU) and natural language generation (NLG). T5 LLM can be understood as a pure transformer using both encoder and decoder networks. It approaches every task as a conversion or generation task, treating them as sequences to sequences. This methodology extends beyond text-to-text tasks to encompass multimodal challenges like text-to-image or image-to-text conversions. For example, in text classification, the encoder accepts text input, while the decoder produces text labels instead of directly classifying them. Encoder-decoder or seq2seq models are commonly employed for tasks demanding both comprehension and generation, where information must be transformed from one format to another, like in machine translation.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pIg7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pIg7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png 424w, https://substackcdn.com/image/fetch/$s_!pIg7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png 848w, https://substackcdn.com/image/fetch/$s_!pIg7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png 1272w, https://substackcdn.com/image/fetch/$s_!pIg7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pIg7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png" width="515" height="341" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2c67842e-d500-4492-aa1e-ab3554486636_515x341.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:341,&quot;width&quot;:515,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pIg7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png 424w, https://substackcdn.com/image/fetch/$s_!pIg7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png 848w, https://substackcdn.com/image/fetch/$s_!pIg7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png 1272w, https://substackcdn.com/image/fetch/$s_!pIg7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c67842e-d500-4492-aa1e-ab3554486636_515x341.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Strengths</h4><p>1. Versatility: Encoder-decoder models like BART and T5 demonstrate remarkable versatility, capable of addressing a wide range of natural language processing tasks, including but not limited to translation, summarization, question answering, and text generation. Their ability to handle diverse tasks within a unified framework simplifies model development and deployment.</p><p>2. Transfer Learning: These models excel at transfer learning, leveraging pretraining on large-scale datasets to acquire generalized language understanding and then fine-tuning on task-specific data to achieve state-of-the-art performance on various NLP tasks. This pretraining-fine tuning paradigm significantly reduces the need for large task-specific datasets and facilitates rapid development and deployment of NLP systems.</p><p>3. Bidirectionality and Contextual Understanding: Models like BART leverage bidirectional architectures, allowing them to capture contextual information from both past and future tokens. This bidirectional understanding enhances their ability to comprehend and generate coherent and contextually appropriate text, improving performance across a wide range of tasks.</p><p>4. Unified Text-to-Text Framework: T5's text-to-text approach simplifies model training and deployment by framing all tasks as text transformations. This unified framework eliminates the need for task-specific architectures or fine-tuning strategies, streamlining the development process and enabling seamless adaptation to new tasks.</p><h4>Limitations</h4><p>1.Computational Resources: Training and deploying encoder-decoder models like BART and T5 require significant computational resources, including high-performance GPUs or TPUs and large-scale datasets. This computational overhead may pose challenges for researchers or organizations with limited resources, hindering widespread adoption and deployment.</p><p>2.Interpretability: The complex nature of encoder-decoder models can make them less interpretable compared to simpler models. Understanding how these models arrive at their predictions, particularly for complex tasks like text generation, may be challenging, limiting their applicability in sensitive domains where interpretability is critical.</p><p>3.Data Efficiency: While encoder-decoder models excel at transfer learning, achieving optimal performance often requires large-scale pretraining datasets and task-specific fine-tuning datasets. In scenarios where labeled data is scarce or of low quality, the benefits of transfer learning may be limited, and models may struggle to generalize to new tasks or domains.</p><p>4.Robustness to Adversarial Inputs: Encoder-decoder models, like other deep learning models, may be susceptible to adversarial attacks, where small perturbations to input data can lead to incorrect or unintended outputs. Ensuring the robustness of these models to adversarial inputs remains an ongoing challenge in NLP research.</p><h4>Notable encoder-decoder models</h4><p>Examples of encoder-decoder models include T5, BART, and BigBird. These models excel in capturing the intricate relationships between inputs and outputs across diverse tasks, enabling versatile applications in natural language processing and beyond.</p><p>Overview of unified LM pre-training. The model parameters are shared across the LM objectives (i.e., bidirectional LM, unidirectional LM, and sequence-to-sequence LM). Courtesy of [Unified Language Model Pre-training arxiv.1905.03197].</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A8ll!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A8ll!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png 424w, https://substackcdn.com/image/fetch/$s_!A8ll!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png 848w, https://substackcdn.com/image/fetch/$s_!A8ll!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png 1272w, https://substackcdn.com/image/fetch/$s_!A8ll!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A8ll!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png" width="849" height="600" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:600,&quot;width&quot;:849,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A8ll!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png 424w, https://substackcdn.com/image/fetch/$s_!A8ll!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png 848w, https://substackcdn.com/image/fetch/$s_!A8ll!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png 1272w, https://substackcdn.com/image/fetch/$s_!A8ll!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e0b1e77-5dd0-43d7-89de-e8f9c893be7a_849x600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Key Technical Concepts in LLMs</h2><p>In this section, we'll explore two fundamental pillars: self-attention and pre-training objectives and strategies.</p><h3>Self-attention</h3><p>Decoding Relationships within Words.Imagine reading a sentence, not just word by word, but understanding how each word relates to the others. That's the essence of self-attention, a mechanism that empowers LLMs.It can capture dependencies and relationships within input sequences. It allows the model to identify and weigh the importance of different parts of the input sequence by attending to itself.Self-attention operates by transforming the input sequence into three vectors: query, key, and value. These vectors are obtained through linear transformations of the input. The attention mechanism calculates a weighted sum of the values based on the similarity between the query and key vectors. The resulting weighted sum, along with the original input, is then passed through a feed-forward neural network to produce the final output. This process allows the model to focus on relevant information and capture long-range dependencies.</p><p>Think of it like attending a party:</p><blockquote><p><em>Guests: Words in the sentence.</em></p><p><em>Conversations: Connections and relationships between words.</em></p><p><em>Attention: Focusing on specific conversations (word relationships) relevant to understanding the overall meaning.</em></p></blockquote><p>Self-attention allows LLMs to:</p><p>Identify crucial connections:: It analyzes all word pairs simultaneously, understanding how each word influences and interacts with others in the sentence.</p><p>Contextual awareness: By focusing on relevant relationships, LLMs gain a deeper understanding of the sentence's meaning, considering the context of each word.</p><p>Long-range dependencies: Unlike traditional models that struggle with long sentences, self-attention can capture relationships between words even if they're far apart in the sentence.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!V1Fs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!V1Fs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png 424w, https://substackcdn.com/image/fetch/$s_!V1Fs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png 848w, https://substackcdn.com/image/fetch/$s_!V1Fs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png 1272w, https://substackcdn.com/image/fetch/$s_!V1Fs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!V1Fs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png" width="816" height="892" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:892,&quot;width&quot;:816,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:55707,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!V1Fs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png 424w, https://substackcdn.com/image/fetch/$s_!V1Fs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png 848w, https://substackcdn.com/image/fetch/$s_!V1Fs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png 1272w, https://substackcdn.com/image/fetch/$s_!V1Fs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f8a053a-917f-4b51-9181-8d9ac7686714_816x892.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Self-attention in this sentence captures the relationship between "it" and "the fish," understanding that "it" refers back to the subject "the dog.</p><h3>Pre-training: Learning Before the Task Begins</h3><p>"Imagine sending a child to school without any prior knowledge. That's how traditional language models were trained &#8211; starting from scratch for each specific task. However, LLMs leverage a powerful technique called pre-training to gain foundational knowledge before tackling specific tasks.</p><p>Think of it like preparing for a job interview:</p><p>General knowledge building: LLMs are exposed to massive amounts of text data (books, articles, code), learning the building blocks of language and general relationships between words.</p><p>Adaptability: This pre-trained knowledge acts as a foundation, allowing LLMs to adapt and learn new tasks more efficiently.</p><p>Fine-tuning: Once the foundation is laid, LLMs are fine-tuned on specific tasks using smaller datasets relevant to that task. This refines their skills for targeted performance.</p><p>Below figure shows the pretraining and fine tuning BERT&nbsp;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UjuZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UjuZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png 424w, https://substackcdn.com/image/fetch/$s_!UjuZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png 848w, https://substackcdn.com/image/fetch/$s_!UjuZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png 1272w, https://substackcdn.com/image/fetch/$s_!UjuZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UjuZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png" width="1456" height="580" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:580,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UjuZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png 424w, https://substackcdn.com/image/fetch/$s_!UjuZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png 848w, https://substackcdn.com/image/fetch/$s_!UjuZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png 1272w, https://substackcdn.com/image/fetch/$s_!UjuZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c7d700-cf72-4c8f-9f1d-8788e64fe9f7_1600x637.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Pre-training objectives and strategies are pivotal in shaping the capabilities and performance of LLMs. Here are some key concepts in this domain:</p><p>Masked Language Modeling (MLM): MLM involves randomly masking a percentage of tokens in the input sequence and training the model to predict the original tokens based on the context provided by the unmasked tokens. This objective, used in models like BERT, facilitates bidirectional learning and enables the model to capture contextual information effectively.</p><p>Autoregressive Language Modeling (ALM): ALM requires the model to predict the next token in a sequence given the preceding tokens. Models like GPT use ALM during pre-training, allowing them to generate coherent and contextually appropriate text.</p><p>Multi-Task Learning: Multi-task learning involves jointly training a model on multiple related tasks during pre-training. This approach enhances the model's ability to learn robust representations by leveraging the shared information across tasks.</p><p>Large-Scale Pre-Training Data: Pre-training LLMs on large-scale text corpora, such as Wikipedia articles or Common Crawl data, helps capture diverse linguistic patterns and semantics, leading to better generalization to downstream tasks.</p><p>Fine-Tuning and Transfer Learning: After pre-training, LLMs are fine-tuned on task-specific data or downstream tasks to adapt their representations to the specific characteristics of the task. Transfer learning from pre-trained LLMs has become a standard practice in NLP, enabling efficient development and deployment of models for various applications.</p><p>Below figure shows a high-level overview of GPT pretraining, and fine-tuning steps. Courtesy of OpenAI.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nOT2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nOT2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png 424w, https://substackcdn.com/image/fetch/$s_!nOT2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png 848w, https://substackcdn.com/image/fetch/$s_!nOT2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png 1272w, https://substackcdn.com/image/fetch/$s_!nOT2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nOT2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png" width="1456" height="623" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:623,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nOT2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png 424w, https://substackcdn.com/image/fetch/$s_!nOT2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png 848w, https://substackcdn.com/image/fetch/$s_!nOT2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png 1272w, https://substackcdn.com/image/fetch/$s_!nOT2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a7747c9-2de9-4fa4-9176-ad1e2821f4f4_1528x654.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Applications of LLMs</h2><p>In this chapter, we explore some of the key applications of LLMs, highlighting their capabilities and impact on tasks such as text generation, question answering, summarization, and translation.</p><p>Below Chart shows the capabilities of LLM [LLM Survey arxiv.2402.06196]</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qNsn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qNsn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png 424w, https://substackcdn.com/image/fetch/$s_!qNsn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png 848w, https://substackcdn.com/image/fetch/$s_!qNsn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png 1272w, https://substackcdn.com/image/fetch/$s_!qNsn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qNsn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png" width="1456" height="635" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:635,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qNsn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png 424w, https://substackcdn.com/image/fetch/$s_!qNsn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png 848w, https://substackcdn.com/image/fetch/$s_!qNsn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png 1272w, https://substackcdn.com/image/fetch/$s_!qNsn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9fc8bd03-8e5f-45a0-8b90-dc3d862363cf_1600x698.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Text Generation</h3><p>Text generation is one of the core capabilities of LLMs, allowing them to produce coherent and contextually relevant text based on given prompts or input sequences. By leveraging the contextual information encoded in their parameters, LLMs can generate text that exhibits characteristics similar to human-written language. This capability has numerous applications in various domains, including content creation, dialogue systems, and creative writing.</p><p>For example, LLMs like OpenAI's GPT series , Google&#8217;s Gemini Ultra have demonstrated remarkable proficiency in generating realistic and diverse text across a wide range of topics and styles. These models can be fine-tuned on specific datasets or prompts to tailor their output to particular domains or tasks, making them versatile tools for content generation.</p><h4>Question Answering</h4><p>LLMs excel at question answering tasks, where they are tasked with generating accurate and informative responses to user queries based on given context or knowledge sources. By leveraging their understanding of language semantics and context, LLMs can effectively extract and synthesize relevant information from large corpora or knowledge bases to answer questions across various domains.</p><p>let's consider a scenario where we want to generate a continuation of a given text prompt. We'll use the GPT-4 model as an example, which is capable of generating thousands of coherent words 2.</p><p>First, we need to prepare our environment and load the pre-trained model. Assuming we have a Python environment set up with the necessary libraries installed, we would load the model and its corresponding tokenizer using the Hugging Face Transformers library.We define a function to generate text based on a given prompt.&nbsp;</p><p><code>from transformers import AutoTokenizer, AutoModelForCausalLM</code></p><p><code># Load Mixtral model and tokenizer</code></p><p><code>model_id = "mistralai/Mixtral-8x7B-Instruct-v0.1"</code></p><p><code>tokenizer = AutoTokenizer.from_pretrained(model_id)</code></p><p><code>model = AutoModelForCausalLM.from_pretrained(model_id)</code></p><p><code># Define a function to generate text using Mixtral</code></p><p><code>def generate_text(prompt, model, tokenizer):</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;inputs = tokenizer(prompt, return_tensors="pt")</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;output = model.generate(inputs, max_length=100, num_return_sequences=1)</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;return tokenizer.decode(output[0], skip_special_tokens=True)</code></p><p><code># Use the function to generate text</code></p><p><code>prompt = "The quick brown fox jumps over the lazy dog."</code></p><p><code>generated_continuation = generate_text(prompt, model, tokenizer)</code></p><p><code>print(generated_continuation)</code></p><p>This code snippet will print out a continuation of the given text prompt using the Mixtral-8x7B model. Note that the actual text generated will depend on the specific model and its training data.Additionally, to fine-tune the LLM for specific tasks, one might use the Hugging Face TRL, Transformers, and Datasets libraries. The fine-tuning process involves setting up training arguments, loading a pre-trained model, and then training it on a custom dataset using the SFTTrainer. After training, the model can be evaluated to measure its performance on a validation set.</p><h3>Summarization</h3><p>LLMs play a crucial role in automatic summarization, where they are employed to generate concise and informative summaries of longer documents or text passages. By distilling the essential information from input texts, LLMs enable users to quickly grasp the key points and main ideas without having to read through the entire document.</p><p>Models like Google's Pegasus and Hugging Face's Bart have demonstrated impressive performance on abstractive summarization tasks, where they generate summaries that go beyond mere extraction of sentences and instead produce coherent and contextually relevant summaries. These models leverage their understanding of language semantics and structure to ensure the fluency and coherence of the generated summaries, making them valuable tools for content summarization and information extraction tasks.</p><p><code>#pip install transformers</code></p><p><code>import logging</code></p><p><code>from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, pipeline</code></p><p><code># Function to load the Llama model</code></p><p><code>def load_model(model_id):</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;logging.info(f"Loading Model: {model_id}")</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;tokenizer = AutoTokenizer.from_pretrained(model_id)</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;model = AutoModelForSeq2SeqLM.from_pretrained(model_id)</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;summarization_pipeline = pipeline('summarization', model=model, tokenizer=tokenizer)</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;logging.info("Model loaded successfully.")</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;return summarization_pipeline</code></p><p><code># Function to summarize text</code></p><p><code>def summarize_text(text, summarization_pipeline):</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;summary = summarization_pipeline(text, max_length=150, min_length=40, do_sample=False)</code></p><p><code>&nbsp;&nbsp;&nbsp;&nbsp;return summary[0]['summary_text']</code></p><p><code># Example usage</code></p><p><code>model_id = "meta-llama/Llama-2-7b-hf"" model ID for summarization</code></p><p><code>summarization_pipeline = load_model(model_id)</code></p><p><code>text_to_summarize = """</code></p><p><code>A long document or passage that needs to be summarized goes here. This could be a news article, academic paper, or any lengthy piece of text. The goal is to condense the information into a shorter version that captures the main points without losing the essence of the original content.</code></p><p><code>"""</code></p><p><code>summary = summarize_text(text_to_summarize, summarization_pipeline)</code></p><p><code>print(summary)</code></p><p>You can use the following code to load the Llama model and perform text summarization.The summarize_text function takes the text you want to summarize and the summarization pipeline you've set up. It returns the summarized text.</p><h3>Translation</h3><p>LLMs have transformed the field of machine translation, enabling the development of more accurate and fluent translation systems. By leveraging their multilingual capabilities and knowledge of language semantics, LLMs can translate text between different languages with remarkable accuracy and fluency.</p><p>For example, models like Google's T5 and Facebook's Marian have demonstrated state-of-the-art performance on multilingual translation tasks, outperforming traditional machine translation systems in terms of translation quality and fluency. These models can effectively capture the nuances and idiosyncrasies of different languages, producing translations that are natural-sounding and contextually appropriate.</p><p>According to the LLM Leaderboard Best Models list on Hugging Face, the model alchemonaut/QuartetAnemoi-70B-t0.0001 is currently considered the best base model for translation tasks. To use this model for translation, you would set up your Python environment with the necessary libraries and then load the model along with its tokenizer using the Hugging Face Transformers library.</p><p>Here's an example of how to use the QuartetAnemoi-70B model for translation:</p><p><code>from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, pipeline</code></p><p><code># Set up the model and tokenizer</code></p><p><code>model_name = "alchemonaut/QuartetAnemoi-70B-t0.0001"</code></p><p><code>tokenizer = AutoTokenizer.from_pretrained(model_name)</code></p><p><code>model = AutoModelForSeq2SeqLM.from_pretrained(model_name)</code></p><p><code># Create the translation pipeline</code></p><p><code>translation_pipeline = pipeline('translation_en_to_de', model=model, tokenizer=tokenizer)</code></p><p><code># Translate text from English to German</code></p><p><code>english_text = "Translate this text to German."</code></p><p><code>translated_text = translation_pipeline(english_text)[0]['translation_text']</code></p><p><code>print(translated_text)</code></p><p>In this example, we assume a translation pipeline for English to German ('translation_en_to_de'). You would need to adjust the pipeline string according to the language pair you are working with. The QuartetAnemoi-70B model is selected based on its standing on the leaderboard, which suggests it performs well in translation tasks 1.</p><p>Remember to replace 'translation_en_to_de' with the appropriate pipeline for your desired language pair. Additionally, ensure that you have the necessary permissions and resources to use this model, considering its size and the computational requirements involved in running such a large model for translation tasks.</p><h2>Limitation and Challenges of LLM</h2><p>Indistinguishability between Generated and Human-Written Text: Large Language Models (LLMs) can generate text that is so fluent and coherent that it is often indistinguishable from human-written content. This poses a significant challenge as it becomes difficult to tell if the output is from a human or the model itself. Detecting LLM-generated text is crucial for several reasons, including combating misinformation, preventing plagiarism, and thwarting impersonation and fraud. Approaches to detecting LLM-generated text include visualizing statistically improbable tokens, using energy-based models, and examining authorship attribution problems. Techniques such as watermarking, which involves adding hidden patterns to generated text, have been investigated for their effectiveness in identifying LLM-generated text. However, adversaries can attempt to evade detection by rephrasing the generated text to remove distinctive signatures, a challenge known as paraphrasing attacks.</p><p>Biases and Toxicity: LLMs can inherit biases from their training data, which can lead to biased outputs. This is particularly concerning because if the training data contains toxic content, such as hate speech or discriminatory language, the LLM can also generate such content. Mitigation strategies for these biases include fine-tuning the model with human preferences or instructions and the development of toolkits for debiasing models. It is also important to ensure that the model does not generate inappropriate or harmful content, which can be a significant ethical concern.</p><p>Prompt Injections and Agency: LLMs can be sensitive to prompt injections, which can lead to unsafe behavior. This is because the responses generated by LLMs are based on the prompts given to them, and if the prompts are manipulated, the model can produce outputs that are not aligned with human values or intentions. Additionally, LLMs lack a sense of agency and do not have the ability to understand the implications of their actions, which can lead to misaligned responses. Addressing these challenges requires a combination of technical innovation, ethical considerations, and ongoing research efforts.</p><p>Unfathomable Datasets: The large size of pre-training datasets used for LLMs makes it nearly impossible for individuals to assess the content thoroughly. This issue, referred to as "Unfathomable Datasets," is a significant challenge because the sheer volume of data can lead to problems with data quality, duplication, and contamination. Issues with near-duplicates in the data can harm model performance, and benchmark data contamination arises when the training data overlaps with the evaluation test set, leading to inflated performance metrics. Addressing these challenges is crucial to ensure accurate, diverse, and bias-free data.</p><p>Privacy Concerns: Personally Identifiable Information (PII) has been found within pre-training data, which can cause privacy breaches during prompting. This is a significant concern because the use of PII in LLMs can lead to the disclosure of sensitive information, and it is essential to address this issue to protect the privacy of individuals.</p><p>Contextual Understanding: LLMs can struggle with understanding context, which can lead to incorrect or nonsensical responses. This is because LLMs generate responses based on patterns learned from their training data, and if the context of a prompt is not clearly defined or is complex, the model may not be able to generate an appropriate response. Overcoming this challenge requires the development of models that can better understand the context in which they are operating.</p><p>Generating Misinformation: LLMs can generate content that is factually incorrect or misleading because they generate responses based on patterns learned from their training data. This can lead to the spread of misinformation, which is a significant concern, especially in the context of information dissemination and decision-making. Addressing this challenge requires the development of techniques to detect and mitigate misleading or false outputs.</p><p>Ethical Concerns: LLMs can be used to generate deep fake text or to automate the creation of misleading news articles or propaganda. This raises serious ethical concerns, as the use of LLMs in this manner can be used to manipulate public opinion and spread misinformation. Addressing these ethical concerns requires a combination of technical innovation, ethical considerations, and ongoing research efforts.</p><p>Lack of Creativity: LLMs are pattern recognition systems and do not truly understand or create new content in the same way a human would. This can be a limitation in situations where creativity and novelty are required. Addressing this challenge requires the development of models that can generate content that is not only accurate and relevant but also creative and novel.</p><p>Controllability: Controlling the output of LLMs is a significant challenge. They can sometimes generate content that is inappropriate, biased, or factually incorrect. This can lead to unintended or harmful outcomes, and it is important to develop methods to control the behavior of LLMs and ensure that they generate outputs that are aligned with human values and objectives.</p><p>Quality of Generated Text: The quality of the output from LLMs can sometimes be inconsistent, producing outputs that are nonsensical or even harmful. This can affect applications like content moderation, information dissemination, and decision-making, underscoring the importance of addressing these issues. Strategies such as pre-training with human feedback, efficient fine-tuning methods, watermarking, and robust evaluation techniques help mitigate challenges and enhance the usability of LLMs.</p><p>Adaptability: While LLMs can handle a wide variety of tasks, they are not specifically designed for fine-tuning, which can make them less adaptable to specific tasks. This can be a limitation in situations where a model needs to be adapted to a specific task or domain. Addressing this challenge requires the development of more efficient fine-tuning methods and the exploration of techniques like pre-training domain mixtures and fine-tuning task mixtures.</p><h2>Evaluating LLMs</h2><p>Evaluating Large Language Models (LLMs) is a multifaceted process that goes beyond mere performance metrics. It encompasses the assessment of accuracy, safety, and fairness, ensuring that LLMs are not only effective but also ethically sound and user-friendly.</p><p>- Bias and Ethical Oversight: Evaluations are crucial for spotlighting biases and controversial outputs, thereby encouraging ethical AI practices and fostering unbiased AI solutions. They serve as a check, ensuring that the outputs of LLMs do not propagate harm, misinformation, or biases, and that the models uphold ethical guidelines.</p><p>- Boosting User Experience: Ensuring that AI-generated content aligns with user needs is a key aspect of LLM evaluations. By validating that AI consistently aligns with societal expectations, trust is built, enhancing user-AI engagement.</p><p>- Versatility and Expertise: Assessments of LLMs reveal their breadth across topics, adaptability to various writing styles, and domain proficiency. This is vital for understanding the model's versatility and its ability to perform in different contexts, whether it's in legal jargon, medical terms, or technical writing.</p><p>- Meeting Regulatory Standards: Evaluations ascertain that LLMs meet the prevailing legal and ethical benchmarks. This is essential for ensuring that the models operate within the legal framework and do not infringe on user privacy or data security standards.</p><p>- Spotting Shortcomings: Evaluations help identify weaknesses in LLMs, whether in nuanced understanding or intricate query resolution. This is a critical step in the development process, allowing for improvements and refinements to be made.</p><p>- Real-world Validation: Practical tests are essential for validating the real-world utility of LLMs. These tests prove the worth of LLMs in actual scenarios, beyond controlled environments, which is vital for assessing their effectiveness in real-world applications.</p><p>- Upholding Accountability: Evaluations ensure responsible AI releases and hold creators accountable for their outputs. This is a critical aspect of maintaining trust in AI technologies and ensuring that they are developed and deployed responsibly.</p><p>Evaluation methods for LLMs can be subjective and time-consuming, especially when using human evaluators. Therefore, it's important to combine human evaluation with quantitative metrics for a comprehensive and efficient evaluation process. Human evaluation provides a qualitative depth, but it can be subjective and vary based on individual perspectives and biases. It can also be time-consuming and expensive compared to automated metrics.To address these challenges, Zero-shot Evaluation is used, where the model is tested on tasks it has not been trained on. This method provides a more realistic measure of the model's capabilities in the real world. Additionally, Performance Assessment involves evaluating how well LLMs generate text and respond to input, using metrics such as accuracy, fluency, coherence, and subject relevance.Model Comparison is another important aspect of LLM evaluation, helping researchers and practitioners compare different models and measure progress. This aids in the selection of the most appropriate model for a given application.Benchmarking Steps for evaluating LLM performance include model evaluation, where models are measured based on their ability to generate accurate, coherent, and contextually appropriate responses, and comparative analysis, where the evaluation results are analyzed to compare the performance of different LLM models on each benchmark task.Common Evaluation Methods include perplexity, human evaluation, BLEU, ROUGE, and diversity measures. However, these methods have their limitations, such as subjectivity, high cost of human evaluations, limited reference data, lack of diversity metrics, and generalization to real-world scenarios.Best Practices for assessing LLMs include using diverse datasets, multi-faceted evaluation, real-world testing, regular updates, feedback loops, inclusive evaluation teams, open peer review, continuous learning, scenario-based testing, and ethical considerations. Evaluation frameworks like Big Bench, GLUE Benchmark, and SuperGLUE are essential for measuring and benchmarking model capabilities, helping in pinpointing a model's strengths, weaknesses, and performance across diverse contexts.</p><p>Challenges with Current Evaluation Techniques include inconsistency across evaluations, which might yield different results for the same model, making it difficult to derive a consistent understanding of an LLM's capabilities. In the evolving world of NLP, accurately evaluating LLMs is crucial. Future evaluation methodologies will likely accentuate context, emotional resonance, and linguistic subtleties, with ethical considerations taking center stage.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dWHg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dWHg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png 424w, https://substackcdn.com/image/fetch/$s_!dWHg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png 848w, https://substackcdn.com/image/fetch/$s_!dWHg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png 1272w, https://substackcdn.com/image/fetch/$s_!dWHg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dWHg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png" width="1112" height="1214" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1214,&quot;width&quot;:1112,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:296373,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dWHg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png 424w, https://substackcdn.com/image/fetch/$s_!dWHg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png 848w, https://substackcdn.com/image/fetch/$s_!dWHg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png 1272w, https://substackcdn.com/image/fetch/$s_!dWHg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91bde1ac-b3af-45f5-a0aa-3be5378fc613_1112x1214.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0gZK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0gZK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png 424w, https://substackcdn.com/image/fetch/$s_!0gZK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png 848w, https://substackcdn.com/image/fetch/$s_!0gZK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png 1272w, https://substackcdn.com/image/fetch/$s_!0gZK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0gZK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png" width="1120" height="950" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:950,&quot;width&quot;:1120,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:246997,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0gZK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png 424w, https://substackcdn.com/image/fetch/$s_!0gZK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png 848w, https://substackcdn.com/image/fetch/$s_!0gZK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png 1272w, https://substackcdn.com/image/fetch/$s_!0gZK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5f1d509-d37c-433a-ad9b-bf9de9e28185_1120x950.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Evaluation metrics and framework</h3><p>G-Eval: This framework uses LLMs to evaluate LLM outputs by generating a series of evaluation steps and then using these steps to determine a final score. It requires several pieces of information to work and can generate a reason for its evaluation score.</p><p>GPTScore: This metric uses the conditional probability of generating the target text as an evaluation metric, differing from G-Eval which directly performs the evaluation task.</p><p>Langchain: Langchain provides various types of evaluators, including string evaluators for assessing the accuracy of predicted strings, trajectory evaluators for analyzing decision-making processes, and comparison evaluators for contrasting outcomes of two different runs on the same input.</p><p>LLM-Eval: This method evaluates multiple dimensions of conversation quality using a single LLM prompt, offering a robust solution with a high correlation with human judgments across diverse datasets.</p><p>LLM-as-a-judge: This approach uses LLMs as a surrogate for human evaluation, demonstrating that LLM judges can achieve an agreement rate exceeding 80% with human evaluations.</p><p>Deep-Eval: An open-source evaluation framework that offers a variety of default metrics like Hallucination, Answer Relevance, Bias, Toxicity, and more, and allows for custom metric creation.</p><p>Llama-Index: Offers evaluation tools for RAG applications, including response evaluation to ensure the response is in line with the retrieved context and retrieval evaluation to gauge the relevance of retrieved sources to the original query.</p><p>RAGAS: Incorporates traditional NLP metrics like BERTScore and NLI for well-rounded evaluations and uses self-consistency methods to ensure reliability in evaluation results.</p><p>These metrics and frameworks are designed to assess various aspects of LLM performance, including coherence, relevance, bias, and more, providing a comprehensive evaluation of the model's output.</p><h2>LLM Books</h2><p>In Chapter - Introduction to NLP, we encountered a handful of books. Here's a refined selection for you to jumpstart your LLM journey.</p><p>1. Practical Natural Language Processing - O'Reilly by Sowmya Vajjala, Bodhisattwa Majumder, Anuj Gupta, Harshit Surana - This book is a comprehensive guide to the practical aspects of NLP, including the use of LLMs. It covers a wide range of topics from the basics of NLP to advanced techniques for processing and generating natural language.</p><p>2. Natural Language Processing with Transformers - O'Reilly by Lewis Tunstall, Leandro von Werra, Thomas Wolf - This book dives into the use of transformer models, which are the backbone of many LLMs like GPT and BERT. It provides a detailed exploration of transformer architectures and their applications in NLP tasks.</p><p>3. Transformers for Natural Language Processing - Packt by Denis Rothman - This book focuses on the transformer models that are central to LLMs, offering an in-depth look at how these models can be used for NLP tasks. It covers both the theoretical aspects of transformers and their practical applications.</p><p>4. GPT-3: Building Innovative NLP Products Using Large Language Models - O'Reilly by Sandra Kublik and Shubham Saboo - This book is specifically about GPT-3, one of the most well-known LLMs. It provides insights into how to build NLP products using GPT-3 and offers practical examples and case studies.</p><p>5. Hands-On Generative AI with Transformers and Diffusion Models -O'Reilly by Pedro Cuenca, Apolinario Passos, Omar Sanseviero, Jonathan Whitaker - This book is a hands-on guide to working with generative AI models, including LLMs and diffusion models. It offers practical exercises and examples to help readers get started with these advanced AI techniques.</p><p>6. Quick Start Guide to Large Language Models by Sinan Ozdemir - The Practical, Step-by-Step Guide to Using LLMs at Scale in Projects and Product.Large Language Models (LLMs) like ChatGPT are demonstrating breathtaking capabilities, but their size and complexity have deterred many practitioners from applying them. In Quick Start Guide to Large Language Models, pioneering data scientist and AI entrepreneur Sinan Ozdemir clears away those obstacles and provides a guide to working with, integrating, and deploying LLMs to solve practical problems.</p><p>7.Understanding Large Language Models: Learning Their Underlying Concepts and Technologies by Thimira Amaratunga - This book will teach you the underlying concepts of large language models (LLMs), as well as the technologies associated with them.The book starts with an introduction to the rise of conversational AIs such as ChatGPT, and how they are related to the broader spectrum of large language models. From there, you will learn about natural language processing (NLP), its core concepts, and how it has led to the rise of LLMs. Next, you will gain insight into transformers and how their characteristics, such as self-attention, enhance the capabilities of language modeling, along with the unique capabilities of LLMs. The book concludes with an exploration of the architectures of various LLMs and the opportunities presented by their ever-increasing capabilities&#8212;as well as the dangers of their misuse.</p><p>8. Generative AI with LangChain: Build large language model (LLM) apps with Python, ChatGPT, and other LLMs by Ben Auffarth - Get to grips with the LangChain framework to develop production-ready applications, including agents and personal assistants. Code examples are regularly updated on GitHub to keep you abreast of the latest LangChain developments.With this book you learn how to leverage LLMs&#8217; capabilities and work around their inherent weaknesses Delve into the realm of LLMs with LangChain and go on an in-depth exploration of their fundamentals, ethical dimensions, and application challenges.Get better at using ChatGPT and GPT models, from heuristics and training to scalable deployment, empowering you to transform ideas into reality</p><p>9. Generative AI on AWS - O'Reilly by Chris Fregly, Antje Barth, Shelbee Eigenbrode - Companies today are moving rapidly to integrate generative AI into their products and services. But there's a great deal of hype (and misunderstanding) about the impact and promise of this technology. This book authors from AWS help CTOs, ML practitioners, application developers, business analysts, data engineers, and data scientists find practical ways to use this exciting new technology.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D65r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D65r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png 424w, https://substackcdn.com/image/fetch/$s_!D65r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png 848w, https://substackcdn.com/image/fetch/$s_!D65r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png 1272w, https://substackcdn.com/image/fetch/$s_!D65r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D65r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png" width="994" height="1162" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1162,&quot;width&quot;:994,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:935564,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D65r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png 424w, https://substackcdn.com/image/fetch/$s_!D65r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png 848w, https://substackcdn.com/image/fetch/$s_!D65r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png 1272w, https://substackcdn.com/image/fetch/$s_!D65r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5e3af90-dfde-46e1-823c-dd6aa1796274_994x1162.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>LLM Libraries</h2><p>LangChain - LangChain is a framework for developing applications that use language models. LangChain provides modules for models, prompts, indices, chains, and agents, with the ability to "chain" these modules together.</p><p>LlamaIndex&nbsp; -&nbsp; LlamaIndex (fka GPT Index) is a framewor for augmenting LLMs with private/custom data. It offers data connectors, indices/graphs that are LLM-compatible, and a query interface over the data.</p><p>LMFlow is an extensible toolbox for finetuning large language models, including LLMs. It supports common backbones such as LLaMA and GPT-.</p><p>Alpaca Farm - AlpacaFarm is a framework for research and development of systems that learn from human feedback (such as instruction-following LLMs). It contains code for simulating preference feedback, automated evaluation, and reinforcement learning algorithms such as PPO.</p><p>Flax - Flax is a neural network library for JAX(a library that provides composable differentiation and vectorization operations for the Python ecosystem) that is designed for flexibility.</p><p>GGML - GGML is a tensor framework for enabling large models to be run on commodity software. It provides integer quantization of models, support for different platforms and intrinsics, and web support. Popular libraries based on GGML include llama.cpp and whisper.cpp.</p><p>Hugging Face&nbsp; - Hugging Face is the Github for machine learning. Hugging Face provides the popular Transformers library for working with the Transformer model, as well as as its hub for machine learning models and datasets.</p><p>Lamini&nbsp; - Lamini is an LLM platform that allows developers to build and run their own custom, private LLMs. Capabilities include fine-tuning, RLHF, and optimizations so the self-hosted LLM runs efficiently.</p><p>MLC LLM&nbsp; - MLC LLM allows language models to be optimized and deployed natively on a broad set of hardware and native applications. Supported platforms include iOS, Android, Apple Silicon, AMD, Intel, NVIDIA, and WebGPU.</p><p>GPTCache: A library for creating a semantic cache to store responses from LLM queries</p><p>Haystack: A library for quickly composing applications with LLM Agents, semantic search, question-answering, and more</p><p>LangFlow: An effortless way to experiment and prototype LangChain flows with drag-and-drop components and a chat interface</p><p>LangKit: An out-of-the-box LLM telemetry collection library that extracts features and profiles prompts, responses, and metadata about how your LLM is performing over time</p><p>LiteLLM: A simple and lightweight package to standardize LLM API calls across various API endpoints</p><p>LLMApp: A Python library that helps you build real-time LLM-enabled data pipelines with few lines of code</p><p>LLMFlows: A framework for building simple, explicit, and transparent LLM applications such as chatbots, question-answering systems, and agents</p><p>LLMonitor: A library for observability and monitoring for AI apps and agents, offering powerful tracing and logging, usage analytics, and deep dives into request histories</p><p>Magentic: A library that seamlessly integrates LLMs as Python functions and allows mixing LLM queries and function calling with regular Python code</p><p>Parea AI: A platform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced prompt playground</p><p>OpenAI: A library that provides access to the capabilities of large language models, facilitating tasks such as text generation, translation, and more.</p><p>Cohere: A library that offers access to a range of language models for various natural language processing tasks.</p><p>Pinecone: A library that provides tools for indexing and retrieving information from LLMs ChatOpenAI: A library that facilitates the creation of chatbots and other interactive applications using LLMs</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lpuV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lpuV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png 424w, https://substackcdn.com/image/fetch/$s_!lpuV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png 848w, https://substackcdn.com/image/fetch/$s_!lpuV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png 1272w, https://substackcdn.com/image/fetch/$s_!lpuV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lpuV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png" width="1000" height="718" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:718,&quot;width&quot;:1000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:266528,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lpuV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png 424w, https://substackcdn.com/image/fetch/$s_!lpuV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png 848w, https://substackcdn.com/image/fetch/$s_!lpuV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png 1272w, https://substackcdn.com/image/fetch/$s_!lpuV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F483b8229-0f37-449a-ad84-b0cae0ecbdb0_1000x718.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><h2>LLM CheatSheet</h2><h3>Purpose</h3><p>The purpose of the LLM Cheat Sheet is to provide a quick and easy-to-use reference guide for NLP practitioners. It covers a wide range of topics related to language modeling and provides a high-level overview of the most essential concepts and techniques in the field. These models are designed to process natural language data and perform various tasks, including text generation, translation, sentiment analysis, and more.</p><h3>Key Concepts</h3><p>Some key concepts to understand when working with large language models include:</p><p>Preprocessing: The input data must be preprocessed before training a language model. This involves cleaning the text, tokenizing it into individual words or subwords, and encoding it in a format that can be fed into the model.</p><p>Fine-tuning: Large language models are often trained on large datasets, but they can also be fine-tuned on smaller, domain-specific datasets to improve their performance on specific tasks.</p><p>Generation: Language models can generate text by predicting the next word in a sequence, or by sampling from a distribution of possible words.</p><p>Translation: Language models can be used for machine translation by encoding text in one language and decoding it into another language.</p><p>Sentiment Analysis: Language models can be used for sentiment analysis by predicting the sentiment of a piece of text, such as whether it is positive, negative, or neutral.</p><p>Tools and Libraries: There are many tools and libraries available for working with large language models.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q9IV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q9IV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png 424w, https://substackcdn.com/image/fetch/$s_!q9IV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png 848w, https://substackcdn.com/image/fetch/$s_!q9IV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png 1272w, https://substackcdn.com/image/fetch/$s_!q9IV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q9IV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png" width="1020" height="260" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:260,&quot;width&quot;:1020,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:28783,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!q9IV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png 424w, https://substackcdn.com/image/fetch/$s_!q9IV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png 848w, https://substackcdn.com/image/fetch/$s_!q9IV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png 1272w, https://substackcdn.com/image/fetch/$s_!q9IV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf6bd088-fce4-4aaf-b43c-b3d411ca8df5_1020x260.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EmkU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EmkU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png 424w, https://substackcdn.com/image/fetch/$s_!EmkU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png 848w, https://substackcdn.com/image/fetch/$s_!EmkU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png 1272w, https://substackcdn.com/image/fetch/$s_!EmkU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EmkU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png" width="1456" height="1125" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1125,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EmkU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png 424w, https://substackcdn.com/image/fetch/$s_!EmkU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png 848w, https://substackcdn.com/image/fetch/$s_!EmkU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png 1272w, https://substackcdn.com/image/fetch/$s_!EmkU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff6912319-f00f-419f-ba02-c4a5d76a4a89_1467x1133.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Quiz questions</h2><p>Here is a quiz to assess understanding of the chapter on Large Language Models (LLMs)</p><p>1. What is the primary purpose of large language models (LLMs)?</p><p>&nbsp;&nbsp;&nbsp;A) To understand and generate human-like language</p><p>&nbsp;&nbsp;&nbsp;B) To predict the next word in a sequence</p><p>&nbsp;&nbsp;&nbsp;C) To perform specific tasks like translation or summarization</p><p>&nbsp;&nbsp;&nbsp;D) All of the above</p><p>2. What are autoregressive models in the context of LLMs?</p><p>&nbsp;&nbsp;&nbsp;A) Models that generate fixed-size vector representations of input text</p><p>&nbsp;&nbsp;&nbsp;B) Models that predict the next word in a sequence given the previous words</p><p>&nbsp;&nbsp;&nbsp;C) Models that focus on different parts of the input sequence during computation</p><p>&nbsp;&nbsp;&nbsp;D) Models that use subword algorithms like BPE or WordPiece</p><p>3. What does the self-attention mechanism in transformer architecture allow the model to do?</p><p>&nbsp;&nbsp;&nbsp;A) Predict the next word in a sequence</p><p>&nbsp;&nbsp;&nbsp;B) Focus on different parts of the input sequence during computation</p><p>&nbsp;&nbsp;&nbsp;C) Analyze its own performance and make adjustments</p><p>&nbsp;&nbsp;&nbsp;D) Improve the generalization capabilities of a model by training it on diverse datasets</p><p>4. What are the key technical concepts in LLMs related to transformer architecture?</p><p>&nbsp;&nbsp;&nbsp;A) Self-attention</p><p>&nbsp;&nbsp;&nbsp;B) Pre-training objectives and strategies</p><p>&nbsp;&nbsp;&nbsp;C) Both A and B</p><p>&nbsp;&nbsp;&nbsp;D) Neither A nor B</p><p>5. Which of the following is an application of large language models?</p><p>&nbsp;&nbsp;&nbsp;A) Sentiment analysis</p><p>&nbsp;&nbsp;&nbsp;B) Question answering</p><p>&nbsp;&nbsp;&nbsp;C) Automatic summarization and Machine Translation</p><p>&nbsp;&nbsp;&nbsp;D) All of the above</p><p>6. What are the pre-training objectives and strategies in LLMs?</p><p>&nbsp;&nbsp;&nbsp;A) Training on a large dataset, usually unsupervised or self-supervised, before fine-tuning for a specific task</p><p>&nbsp;&nbsp;&nbsp;B) Training on a small dataset, usually supervised, for a specific task</p><p>&nbsp;&nbsp;&nbsp;C) Training on a large dataset, usually supervised, for a specific task</p><p>&nbsp;&nbsp;&nbsp;D) Training on a small dataset, usually unsupervised or self-supervised, before fine-tuning for a specific task</p><p>7. What is the main characteristic of autoregressive language models like GPT-3?</p><p>&nbsp;&nbsp;&nbsp;A) They generate fixed-size vector representations of input text</p><p>&nbsp;&nbsp;&nbsp;B) They predict the next word in a sequence given the previous words</p><p>&nbsp;&nbsp;&nbsp;C) They focus on different parts of the input sequence during computation</p><p>&nbsp;&nbsp;&nbsp;D) They use subword algorithms like BPE or WordPiece</p><p>8. What is the main goal of an autoencoding language model like BERT?</p><p>&nbsp;&nbsp;&nbsp;A) To generate fixed-size vector representations of input text</p><p>&nbsp;&nbsp;&nbsp;B) To predict the next word in a sequence given the previous words</p><p>&nbsp;&nbsp;&nbsp;C) To focus on different parts of the input sequence during computation</p><p>&nbsp;&nbsp;&nbsp;D) To use subword algorithms like BPE or WordPiece</p><p>9. Which LLM is an example of a combination of autoencoding and autoregressive language models?</p><p>&nbsp;&nbsp;&nbsp;A) Turing NLG</p><p>&nbsp;&nbsp;&nbsp;B) GPT series</p><p>&nbsp;&nbsp;&nbsp;C) BERT</p><p>&nbsp;&nbsp;&nbsp;D) T5</p><p>10. What does the term 'temperature' refer to in the context of LLMs?</p><p>&nbsp;&nbsp;&nbsp;&nbsp;A) A parameter that affects the randomness of the output of the softmax layer</p><p>&nbsp;&nbsp;&nbsp;&nbsp;B) A measure of the size of the model</p><p>&nbsp;&nbsp;&nbsp;&nbsp;C) A parameter that controls the length of the generated text</p><p>&nbsp;&nbsp;&nbsp;&nbsp;D) A parameter that controls the complexity of the model's predictions</p><p>11. What is the faithfulness metric in RAG evaluation used for?</p><p>A) To measure the length of the LLM output</p><p>B) To evaluate whether the LLM output factually aligns with the retrieval context</p><p>C) To assess the speed of the LLM application</p><p>D) To measure the complexity of the LLM output</p><p>12. Which of the following LLM evaluation metrics is not purely statistical and takes into account semantics and reasoning capabilities?</p><p>A) BLEU</p><p>B) QAG Score</p><p>C) GPTScore</p><p>D) Levenshtein distance</p><p>Correct Answers:</p><p>1.&nbsp; D</p><p>2.&nbsp; B</p><p>3.&nbsp; B</p><p>4.&nbsp; C</p><p>5.&nbsp; D</p><p>6.&nbsp; D</p><p>7.&nbsp; B</p><p>8.&nbsp; A</p><p>9.&nbsp; C</p><p>10. A</p><p>11. B</p><p>12. C</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Abonia Sojasingarayar! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Chapter 1 - Introduction to NLP]]></title><description><![CDATA[Understanding of NLP and its potential applications]]></description><link>https://aboniasojasingarayar.substack.com/p/introduction-to-nlp</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/introduction-to-nlp</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Tue, 23 Jul 2024 06:41:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!LsAa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Natural language processing (NLP) is a field of computer science that deals with the interaction between computers and human language. NLP research has made significant progress in recent years, and it is now possible to develop systems that can understand, generate, translate, and even write human-quality text. It has become a hot topic in AI research because of its many potential applications, such as text generation, chatbots, and text-to-image applications.Recent advancements in NLP have led to a revolution in the ability of computers to understand human languages. These advancements have also allowed computers to understand programming languages and even biological and chemical sequences that resemble language. The latest NLP models are now able to analyze the meanings of input text and generate meaningful, expressive output. This means that computers can now understand and respond to human language in a more natural way.</p><p>In this article, we will provide an introduction to NLP and its applications. We will discuss the basic concepts of NLP, such as tokenization, stemming, and tagging. We will also introduce some of the most common NLP tasks, such as machine translation, sentiment analysis, and text summarization.</p><p>The goal of this article is to give you a basic understanding of NLP and its potential applications. By the end of this article, you will be able to identify the different components of an NLP system and understand the challenges involved in developing these systems.</p><p>In this article, we will cover the following topics,:</p><ul><li><p>What is NLP? Definition and scope</p></li><li><p>Why does NLP matter?&nbsp;</p></li><li><p>Applications of NLP and Historical Context: Real-world uses, Evolution and milestones</p></li><li><p>How does NLP work?</p></li><li><p>Challenges and Limitations of NLP: Technical hurdles and complexities</p></li><li><p>Further Resource : popular NLP datasets, libraries, and online courses</p></li></ul><div><hr></div><h2>What is NLP? Definition and scope</h2><p>Natural language processing (NLP) is a branch of artificial intelligence that deals with the interaction between computers and human (natural) languages. As a branch of artificial intelligence, NLP&nbsp; uses machine learning to process and interpret text and data. Natural language recognition and natural language generation are types of NLP.It encompasses a wide range of tasks, including understanding and generating human language, translating languages, and summarizing text. NLP has become increasingly important in recent years as the amount of data available in the form of text and speech has grown exponentially. NLP has existed for more than 50 years and has roots in the field of linguistics. It has a variety of real-world applications in a number of fields, including medical research, search engines and business intelligence.</p><div><hr></div><h2>What is natural language processing used for?</h2><p>Natural language processing applications are used to derive insights from unstructured text-based data and give you access to extracted information to generate new understanding of that data. Natural language processing examples can be built using Python, TensorFlow, and PyTorch.</p><p>NLP serves various essential roles across diverse industries. In the realm of customer sentiment analysis, NLP employs entity analysis to identify and label specific fields within documents and communication channels, offering valuable insights into customer opinions and facilitating the discovery of product and User Experience (UX) enhancements. For receipt and invoice understanding, NLP extracts entities, such as dates and prices, enabling the recognition of patterns between requests and payments. Document analysis benefits from custom entity extraction, streamlining the identification of domain-specific entities within documents without the need for labor-intensive manual analysis. Content classification involves categorizing documents by common entities, customized domain-specific entities, or over 700 general categories, aiding in tasks like trend spotting and marketing content extraction. In healthcare, NLP enhances clinical documentation, supports data mining research, and expedites automated registry reporting, thereby contributing to the acceleration of clinical trials. In finance, it automates tasks like customer service chatbots, fraud detection, and sentiment analysis of financial news. Legal applications involve analyzing legal documents, summarizing case law, and generating legal contracts. In education, NLP creates personalized learning experiences, assesses student writing, and generates adaptive learning materials. For government use, NLP analyzes public opinion, summarizes government reports, and automates administrative tasks, showcasing its versatility and significance across a spectrum of industries.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LsAa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LsAa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!LsAa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!LsAa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!LsAa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LsAa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png" width="1080" height="1080" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1080,&quot;width&quot;:1080,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!LsAa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png 424w, https://substackcdn.com/image/fetch/$s_!LsAa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png 848w, https://substackcdn.com/image/fetch/$s_!LsAa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!LsAa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd005eec4-0260-4f83-ba87-d48dbb595388_1080x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>Why Does Natural Language Processing (NLP) Matter?</h2><p>Natural language processing is a crucial part of everyday life, and it's becoming even more important as language technology is being applied to a wide range of industries, including healthcare and retail. Conversational agents like Amazon's Alexa and Apple's Siri rely on NLP to understand user inquiries and provide answers. GPT-4, a sophisticated NLP system, is capable of generating high-quality text on various topics and powering chatbots that can hold meaningful conversations. Google utilizes NLP to enhance its search engine results, while social media platforms like Facebook employ it to detect and block hate speech. While NLP is becoming increasingly sophisticated, there's still much room for improvement. Current systems can be biased and inconsistent, and they sometimes behave erratically. However, machine learning engineers have numerous opportunities to apply NLP in ways that will become increasingly essential to society.</p><p>NLP encompasses a broad range of subfields, each of which focuses on a particular aspect of language processing. Some of the most important subfields of NLP include:</p><p>Natural language understanding (NLU): This subfield focuses on the ability of computers to understand the meaning of human language. This includes tasks such as parsing, which involves breaking down text into its constituent parts (e.g., words, phrases, sentences), and semantic analysis, which involves extracting the meaning of the text.</p><p>Natural language generation (NLG): This subfield focuses on the ability of computers to generate human-like text. This includes tasks such as machine translation, which involves translating text from one language to another, and text summarization, which involves generating a concise overview of a piece of text.</p><p>Speech recognition: This subfield focuses on the ability of computers to understand human speech. This includes tasks such as transcribing audio recordings into text and controlling devices with voice commands.</p><p>Text-to-speech synthesis: This subfield focuses on the ability of computers to generate human-like speech. This includes tasks such as converting text into audio recordings and creating avatars that can speak.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cexd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cexd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png 424w, https://substackcdn.com/image/fetch/$s_!cexd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png 848w, https://substackcdn.com/image/fetch/$s_!cexd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png 1272w, https://substackcdn.com/image/fetch/$s_!cexd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cexd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png" width="1000" height="648" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:648,&quot;width&quot;:1000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!cexd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png 424w, https://substackcdn.com/image/fetch/$s_!cexd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png 848w, https://substackcdn.com/image/fetch/$s_!cexd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png 1272w, https://substackcdn.com/image/fetch/$s_!cexd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3b61cb5-c74d-4e0f-87d4-da91ab4e2ab6_1000x648.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Information Extraction: This field is concerned with the extraction of semantic information from text, which includes tasks such as named-entity recognition, coreference resolution, and relationship extraction.</p><p>Ontology Engineering: This field studies the methods and methodologies for building ontologies, which are formal representations of a set of concepts within a domain and the relationships between those concepts.</p><p>Statistical Natural Language Processing: This subfield includes statistical semantics, which establishes semantic relations between words to examine their contexts, and distributional semantics, which examines the semantic relationship of words across a corpora or in large samples of data.</p><div><hr></div><h2>Historical Context and Applications of NLP: Evolution , Milestones and Real-world uses</h2><h3>History and Milestones</h3><p>The early <strong>history </strong>of NLP is marked by the work of Swiss linguist Ferdinand de Saussure, who described language as a system of relationships. In 1950, Alan Turing proposed the Turing test as a way to measure a machine&#8217;s ability to exhibit intelligent behavior equivalent to, or indistinguishable from, that of a human. This led to the development of NLP as a field of study.</p><p>The 1960s saw the development of early NLP systems, including ELIZA, which could simulate a conversation with a human. However, these systems were limited in their ability to understand and generate natural language. In the 1980s, there was a shift in NLP towards the use of machine learning algorithms. This led to significant improvements in the accuracy of NLP systems.</p><p>In the 1990s, the development of statistical NLP models further enhanced the field. These models were able to handle large amounts of text data and make more nuanced linguistic distinctions. In the 2000s, there was a resurgence of interest in NLP, driven by the availability of large amounts of data and advances in computing power. This led to the development of neural network models, which are now the state-of-the-art in NLP.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YW1z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YW1z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png 424w, https://substackcdn.com/image/fetch/$s_!YW1z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png 848w, https://substackcdn.com/image/fetch/$s_!YW1z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!YW1z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YW1z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png" width="640" height="1600" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1600,&quot;width&quot;:640,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!YW1z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png 424w, https://substackcdn.com/image/fetch/$s_!YW1z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png 848w, https://substackcdn.com/image/fetch/$s_!YW1z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png 1272w, https://substackcdn.com/image/fetch/$s_!YW1z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1bf7350-415a-419b-bde2-9318b8c8404c_640x1600.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>As in the above diagram, NLP has a rich history that can be traced back to the early days of computer science. Some of the key <strong>milestones </strong>in the history of NLP include:</p><p>1950: Alan Turing publishes his seminal paper, "Computing Machinery and Intelligence," which introduces the Turing test, a benchmark for evaluating machine intelligence.</p><p>1957: George Miller published his paper, "The Magical Number Seven, Plus or Minus Two: Some Limits on Our Capacity for Processing Information," which highlights the limitations of human short-term memory and has implications for NLP research.</p><p>1960s: The development of statistical methods and machine learning algorithms leads to significant advances in NLP.</p><p>1980s: The rise of personal computers and the availability of large datasets fuel further progress in NLP.</p><p>2000s: The development of deep learning techniques revolutionizes NLP, enabling computers to achieve human-level performance in many tasks.</p><p>2020s: Large language Models, language model notable for its ability to achieve general-purpose language generation. LLMs acquire these abilities by learning statistical relationships from text documents during a computationally intensive self-supervised and semi-supervised training process.&nbsp;</p><div><hr></div><h3>Applications Area</h3><p>Search engines: NLP is used to understand the queries that users enter into search engines and to return relevant results.</p><p>Virtual assistants: NLP is used to power virtual assistants such as Siri, Alexa, and Google Assistant. These assistants can understand natural language commands and respond in a way that is helpful and informative.</p><p>Machine translation: NLP is used to translate text from one language to another. This is a valuable tool for businesses that operate internationally and for people who want to read or communicate with people who speak different languages.</p><p>Sentiment analysis: NLP is used to analyze the sentiment of text, which is the attitude or opinion that the author expresses. This is a valuable tool for businesses that want to understand their customers' feedback and for social media analysts who want to track trends in public opinion.</p><p>Text summarization: Text summarization is the task of condensing lengthy documents into shorter, concise summaries that capture the key points of the original text. This is often done by identifying the most important sentences and phrases, and then rephrasing them in a more concise way. Text summarization is a valuable tool for quickly understanding the content of long documents, and it is used in a variety of applications, including news articles, research papers, and business reports.</p><p>Named entity recognition (NER): Named entity recognition is the task of identifying and extracting important entities like people, places, organizations, and dates from text. This is important for a variety of applications, such as search engines, social media analysis, and machine translation. For example, a search engine could use NER to identify the names of people, places, and organizations in user queries, and then provide more relevant search results.</p><p>Speech recognition: Speech recognition is the task of converting spoken language into text. This is a critical technology for a variety of applications, including voice assistants, dictation software, and call centers. For example, a voice assistant could use speech recognition to understand user commands, and then perform the desired actions.</p><p>Chatbots: Chatbots are computer programs that simulate conversation with human users. They are typically used to provide customer service, answer questions, and provide information. Chatbots are becoming increasingly sophisticated, and they are now able to engage in natural conversations with humans.</p><p>Question answering (QA): Question answering is the task of answering user questions based on text or knowledge bases. This is a valuable tool for accessing information quickly and easily. For example, a QA system could be used to answer questions about a book, a website, or a database.</p><p>Information retrieval (IR): Information retrieval is the task of finding relevant information from large amounts of text. This is a critical task for search engines, libraries, and other information systems. IR systems use a variety of techniques, such as indexing, ranking, and relevance feedback, to identify the most relevant documents for a given query.</p><p>Text generation: Text generation is the task of creating human-like text, such as poems, code, scripts, and musical pieces. This is a complex task that requires a deep understanding of language and the ability to generate creative text formats. Text generation is used in a variety of applications, such as machine translation, chatbots, and creative writing tools.</p><p>Linguistic analysis: Linguistic analysis is the study of the structure and meaning of human language. This includes the study of syntax, semantics, and pragmatics. Linguistic analysis is used to develop natural language processing (NLP) systems, and it is also used to study the evolution of language and the relationship between language and thought.</p><p>In addition to the above applications, NLP is also used in a variety of other fields, such as:</p><blockquote><p>Law:Used to analyze legal documents, identify relevant information, and extract legal concepts.</p><p>Finance:Used to analyze financial data, identify trends, and make predictions.</p><p>Marketing: Used to analyze customer data, personalize marketing messages, and understand customer sentiment.</p><div><hr></div></blockquote><h2>How Does Natural Language Processing (NLP) Work?</h2><p>NLP models work by finding relationships between the constituent parts of language &#8212; for example, the letters, words, and sentences found in a text dataset. NLP architectures use various methods for data preprocessing, feature extraction, and modeling. Some of these processes are:&nbsp;</p><h4>Data Preprocessing</h4><p>Before a model deals with text for a particular task, the text usually needs some preparation. This is called data preprocessing, and it's important for the model to work well. Data-centric AI, a growing approach, focuses on this step. Various techniques are used in data preprocessing:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kZeM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kZeM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png 424w, https://substackcdn.com/image/fetch/$s_!kZeM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png 848w, https://substackcdn.com/image/fetch/$s_!kZeM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png 1272w, https://substackcdn.com/image/fetch/$s_!kZeM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kZeM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png" width="1108" height="1084" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1084,&quot;width&quot;:1108,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:81511,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!kZeM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png 424w, https://substackcdn.com/image/fetch/$s_!kZeM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png 848w, https://substackcdn.com/image/fetch/$s_!kZeM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png 1272w, https://substackcdn.com/image/fetch/$s_!kZeM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4e5f5383-947e-47e0-8c6d-b63ebcbe4746_1108x1084.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Stemming and Lemmatization: These are ways to convert words to their basic forms. Stemming is a more informal method that uses rules to find base forms, while lemmatization is a formal process that analyzes a word's structure using a dictionary. Libraries like spaCy and NLTK provide tools for stemming and lemmatization.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://substack.com/refer/aboniasojasingarayar?utm_source=substack&amp;utm_context=post&amp;utm_content=146161066&amp;utm_campaign=writer_referral_button&quot;,&quot;text&quot;:&quot;Start a Substack&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Start writing today. Use the button below to create your Substack and connect your publication with Abonia Sojasingarayar</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/p/introduction-to-nlp?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Start a Substack&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://aboniasojasingarayar.substack.com/p/introduction-to-nlp?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Start a Substack</span></a></p></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!t1GM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!t1GM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png 424w, https://substackcdn.com/image/fetch/$s_!t1GM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png 848w, https://substackcdn.com/image/fetch/$s_!t1GM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png 1272w, https://substackcdn.com/image/fetch/$s_!t1GM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!t1GM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png" width="1286" height="386" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:386,&quot;width&quot;:1286,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:125390,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!t1GM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png 424w, https://substackcdn.com/image/fetch/$s_!t1GM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png 848w, https://substackcdn.com/image/fetch/$s_!t1GM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png 1272w, https://substackcdn.com/image/fetch/$s_!t1GM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fff7b6e-d9db-4a8f-ba8a-8a97e13928ce_1286x386.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sNdt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sNdt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png 424w, https://substackcdn.com/image/fetch/$s_!sNdt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png 848w, https://substackcdn.com/image/fetch/$s_!sNdt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png 1272w, https://substackcdn.com/image/fetch/$s_!sNdt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sNdt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png" width="1380" height="604" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:604,&quot;width&quot;:1380,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:141155,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!sNdt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png 424w, https://substackcdn.com/image/fetch/$s_!sNdt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png 848w, https://substackcdn.com/image/fetch/$s_!sNdt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png 1272w, https://substackcdn.com/image/fetch/$s_!sNdt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F886b1a04-5a18-4e78-96a0-f28951175edb_1380x604.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Sentence Segmentation: Breaking a big piece of text into meaningful sentence units. In English, this is often marked by a period, but it can be tricky. For instance, a period might indicate an abbreviation and not necessarily the end of a sentence. This becomes even more challenging in languages like ancient Chinese without a clear sentence-ending marker.</p><p>Stop Word Removal: Removing very common words, like "the," "a," and "an," that don't contribute much information to the text.</p><p>Tokenization: Breaking text into individual words and fragments. The result includes a word index and tokenized text where words are represented as numerical tokens for deep learning methods. Ignoring unimportant tokens can make language models more efficient.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SgHV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SgHV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png 424w, https://substackcdn.com/image/fetch/$s_!SgHV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png 848w, https://substackcdn.com/image/fetch/$s_!SgHV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png 1272w, https://substackcdn.com/image/fetch/$s_!SgHV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SgHV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png" width="1380" height="502" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ebbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:502,&quot;width&quot;:1380,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:151437,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!SgHV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png 424w, https://substackcdn.com/image/fetch/$s_!SgHV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png 848w, https://substackcdn.com/image/fetch/$s_!SgHV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png 1272w, https://substackcdn.com/image/fetch/$s_!SgHV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febbec334-abdf-4ac5-818d-f9ad82b2e226_1380x502.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Feature extraction</h4><p>Most conventional machine-learning techniques work on the features &#8211; generally numbers that describe a document in relation to the corpus that contains it &#8211; created by either Bag-of-Words, TF-IDF, or generic feature engineering such as document length, word polarity, and metadata (for instance, if the text has associated tags or scores). More recent techniques include Word2Vec, GLoVE, and learning the features during the training process of a neural network.</p><p>Feature extraction serves as a vital step in transforming raw text into a format that machine learning algorithms can comprehend. These algorithms rely on numerical representations of language to perform tasks like text classification, sentiment analysis, and machine translation.</p><p>Conventional feature extraction techniques, such as Bag-of-Words (BoW) and TF-IDF, involve counting word frequencies or term weights to represent documents. While effective, these methods overlook the contextual relationships between words. To address this limitation, advanced techniques like Word2Vec and GLoVE have emerged. These methods capture semantic relationships between words by learning word embeddings, which are vector representations that preserve the meaning of words.Let's delve into each technique:</p><p>Bag-of-Words (BoW):</p><p>Imagine a document as a bag filled with words. BoW treats each document as a bag, counting the number of times each unique word appears. This approach provides a simple representation but lacks context.The table presented below illustrates the functioning of Bag of Words (BoW).</p><blockquote><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ngn-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ngn-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png 424w, https://substackcdn.com/image/fetch/$s_!ngn-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png 848w, https://substackcdn.com/image/fetch/$s_!ngn-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png 1272w, https://substackcdn.com/image/fetch/$s_!ngn-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ngn-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png" width="1446" height="336" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/52274f25-da25-4b74-839b-39a326faa295_1446x336.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:336,&quot;width&quot;:1446,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:58845,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!ngn-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png 424w, https://substackcdn.com/image/fetch/$s_!ngn-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png 848w, https://substackcdn.com/image/fetch/$s_!ngn-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png 1272w, https://substackcdn.com/image/fetch/$s_!ngn-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52274f25-da25-4b74-839b-39a326faa295_1446x336.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div></blockquote><p>Term Frequency-Inverse Document Frequency (TF-IDF):</p><p>TF-IDF improves upon BoW by considering word frequency and document rarity. It assigns higher weights to words that appear frequently in a particular document but infrequently across the entire corpus, effectively capturing the importance of words.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HrRK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HrRK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png 424w, https://substackcdn.com/image/fetch/$s_!HrRK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png 848w, https://substackcdn.com/image/fetch/$s_!HrRK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png 1272w, https://substackcdn.com/image/fetch/$s_!HrRK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HrRK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png" width="1294" height="840" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:840,&quot;width&quot;:1294,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:194481,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!HrRK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png 424w, https://substackcdn.com/image/fetch/$s_!HrRK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png 848w, https://substackcdn.com/image/fetch/$s_!HrRK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png 1272w, https://substackcdn.com/image/fetch/$s_!HrRK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc6e808a-05f0-40b3-b3b2-db7c21318f8d_1294x840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Word2Vec:</p><p>In 2013, Word2Vec revolutionized the field of NLP by introducing neural networks to learn word embeddings directly from raw text. It offers two main variants: Skip-gram and Continuous Bag-of-Words (CBOW). Skip-gram predicts surrounding words given a target word, while CBOW predicts the target word from surrounding words. Both methods learn word embeddings that capture semantic relationships between words.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kgwZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kgwZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png 424w, https://substackcdn.com/image/fetch/$s_!kgwZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png 848w, https://substackcdn.com/image/fetch/$s_!kgwZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png 1272w, https://substackcdn.com/image/fetch/$s_!kgwZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kgwZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png" width="1398" height="374" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:374,&quot;width&quot;:1398,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:59302,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!kgwZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png 424w, https://substackcdn.com/image/fetch/$s_!kgwZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png 848w, https://substackcdn.com/image/fetch/$s_!kgwZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png 1272w, https://substackcdn.com/image/fetch/$s_!kgwZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa0d02b7-a7d9-4592-99aa-d8bf505f7ac4_1398x374.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>GLoVE (Global Vectors for Word Representation):</p><p>GLoVE shares similarities with Word2Vec but employs matrix factorization techniques rather than neural learning. It constructs a matrix based on global word-to-word co-occurrence counts, effectively capturing semantic relationships while avoiding the computational complexity of neural networks.</p><h4>Modeling</h4><p>After data is preprocessed, it is fed into an NLP architecture that models the data to accomplish a variety of tasks.</p><p>Numerical features extracted by the techniques described above can be fed into various models depending on the task at hand. For example, for classification, the output from the TF-IDF vectorizer could be provided to logistic regression, naive Bayes, decision trees, or gradient boosted trees. Or, for named entity recognition, we can use hidden Markov models along with n-grams.&nbsp;</p><p>Deep neural networks typically work without using extracted features, although we can still use TF-IDF or Bag-of-Words features as an input.&nbsp;</p><p>Language Models: In very basic terms, the objective of a language model is to predict the next word when given a stream of input words. Probabilistic models that use Markov assumption are one example:</p><blockquote></blockquote><p>Deep learning is also used to create such language models. Deep-learning models take as input a word embedding and, at each time state, return the probability distribution of the next word as the probability for every word in the dictionary. Pre-trained language models learn the structure of a particular language by processing a large corpus, such as Wikipedia. They can then be fine-tuned for a particular task. For instance, BERT has been fine-tuned for tasks ranging from fact-checking to writing headlines.&nbsp;</p><div><hr></div><h2>Challenges and Limitations of NLP: Technical hurdles and complexities</h2><p>NLP is a complex field with a number of challenges, including:</p><ol><li><p>Ambiguity: Human language is often ambiguous and can have multiple meanings. This can make it difficult for computers to understand the meaning of text correctly.</p></li><li><p>Word sense disambiguation: Words can have multiple meanings, and it can be difficult for computers to determine the correct meaning in a particular context.</p></li><li><p>Domain dependence: NLP models are often trained on data from a specific domain, such as news articles or technical documents. This can make it difficult for them to generalize to other domains.</p></li><li><p>Scalability: NLP models can be very large and complex, and it can be difficult to train and deploy them in real-world applications.</p></li></ol><div><hr></div><h3>Complexities</h3><p>Beyond technical hurdles above, NLP also faces a range of complexities that stem from the intricate nature of human language. These complexities require NLP models to handle a variety of linguistic phenomena and adapt to different cultural and social contexts.</p><p>Linguistic Phenomena: Human language is rich with linguistic phenomena that NLP models must be able to handle. These include idioms, sarcasm, figurative language, negation, and ellipsis, which can often be interpreted in multiple ways and require sophisticated linguistic analysis.</p><p>Cultural and Social Context: Understanding the nuances of culture and social context is crucial for accurate NLP applications. For instance, NLP systems used in customer service or marketing need to be aware of cultural sensitivities and adapt their language accordingly.</p><p>Ethical Considerations: The development and use of NLP systems raise ethical concerns around bias, discrimination, and privacy. NLP models should be carefully designed and trained to avoid perpetuating biases and respecting user privacy.</p><div><hr></div><h3>Addressing Challenges and Limitations</h3><p>Overcoming these challenges and limitations requires continuous research and innovation in NLP. Researchers are exploring new techniques and algorithms to tackle ambiguity, handle unstructured data, and generalize to domain-specific language.</p><p>Model Complexity: Deep learning models have shown remarkable progress in NLP, but they often require large amounts of data and computational resources. Researchers are developing more efficient and scalable deep learning architectures to address these limitations.</p><p>Domain Adaptation: NLP models are increasingly being adapted to specific domains using techniques such as domain adaptation and transfer learning. These techniques aim to transfer the knowledge learned from a large general-purpose dataset to a smaller dataset of domain-specific text.</p><p>Human-In-the-Loop Systems: Incorporating human expertise and feedback into NLP systems can help improve accuracy and reduce bias. Human-in-the-loop systems can be used for tasks such as data cleaning, labeling, and error correction.</p><p>Explainable AI: Understanding how NLP models make decisions is critical for building trust and transparency. Explainable AI techniques aim to provide insights into the reasoning behind NLP models' predictions, allowing users to identify potential biases and areas for improvement.</p><div><hr></div><h2>Further Resources: Popular NLP Datasets, Libraries, and Online Courses</h2><h3>NLP and LLM Books</h3><ol><li><p>Hands-On Large Language Models by Jay Alammar and Maarten Grootendorst <strong>: </strong>This book provides a practical guide to working with large language models. It was released in December 2024 and is published by O'Reilly Media, Inc.</p></li><li><p>Foundation of Statistical Natural Language Processing by Christopher D Manning and Hinrich Sch&#252;tze : This book explores the statistical approach to Natural Language Processing and helps users master the mathematics and linguistics needed. It focuses on statistical methods and techniques that have recently become popular.</p></li><li><p>Speech and Language Processing by Dan Jurafsky and James H Martin : This book provides a comprehensive introduction to the field of speech and language processing. It covers a wide range of topics including regular expressions and automata, morphology and finite-state transducers, computational phonology and text-to-speech, probabilistic models of pronunciation and spelling, HMMs and speech recognition, word classes and part-of-speech tagging, context-free grammars for English, parsing with context-free grammar, lexicalized and probabilistic parsing, language and complexity, semantic analysis, lexical semantics, machine translation.</p></li><li><p>Natural Language Understanding by James Allen. This book is another introductory guide to NLP and considered a classic. While it was published in 1994, it&#8217;s highly relevant to today&#8217;s discussions and analytics activities and lauded by generations of NLP researchers and educators. It introduces major techniques and concepts required to build NLP systems, and goes into the background and theory of each without overwhelming readers in technical jargon.</p></li><li><p>Handbook of Natural Language Processing by Nitin Indurkhya and Fred J. Damerau. This comprehensive, modern &#8220;Handbook of Natural Language Processing&#8221; offers tools and techniques for developing and implementing practical NLP in computer systems. There are three sections to the book: classical techniques (including symbolic and empirical approaches), statistical approaches in NLP, and multiple applications&#8212;from information visualization to ontology construction and biomedical text mining. The second edition has a multilingual scope, accommodating European and Asian languages besides English, plus there&#8217;s greater emphasis on statistical approaches. Furthermore, it features a new applications section discussing emerging areas such as sentiment analysis. It&#8217;s a great start to learn how to apply NLP to computer systems.</p></li><li><p>The Handbook of Computational Linguistics and Natural Language Processing by Alexander Clark, Chris Fox, and Shalom Lappin. Similar to the &#8220;Handbook of Natural Language Processing,&#8221; this book includes an overview of concepts, methodologies, and applications in NLP and Computational Linguistics, presented in an accessible, easy-to-understand way. It features an introduction to major theoretical issues and the central engineering applications that NLP work has produced to drive the discipline forward. Theories and applications work hand in hand to show the relationship in language research as noted by top NLP researchers. It&#8217;s a great resource for NLP students and engineers developing NLP applications in labs at software companies.</p></li><li><p>Deep Learning for Coders with fastai and PyTorch by Jeremy Howard and Sylvain Gugger <strong>: </strong>This book teaches deep learning concepts through practical coding exercises using the fastai and PyTorch libraries. It's designed for coders who want to understand and implement deep learning models.</p></li><li><p>Natural Language Processing with Transformers, Revised Edition by Lewis Tunstall, Leandro von Werra, and Thomas Wolf <strong>: </strong>Since their introduction in 2017, transformers have quickly become the dominant architecture for achieving state-of-the-art results in NLP. This book provides a comprehensive guide to implementing transformers in NLP.</p></li><li><p>Practical Natural Language Processing by Timothy Baldwin and Tristan Jehan. This book provides practical advice on how to implement NLP systems in real-world scenarios. It covers a wide range of topics, from basic techniques to advanced methods, and is suitable for both beginners and experienced professionals in the field.</p></li><li><p>Natural Language Processing with Transformers by Thomas Wolf. This book focuses on the use of transformers in NLP, offering a comprehensive introduction to the topic. It covers everything from the basics of transformers to their application in various NLP tasks.&nbsp;</p></li><li><p>Transformers for Natural Language Processing by Abhishek Thakur. This book dives deep into the principles and techniques behind transformers, making it a great resource for those interested in exploring this cutting-edge technology in the field of NLP.</p></li><li><p>&#8220;The Oxford Handbook of Computational Linguistics&#8221; by Ruslan Mitkov .This handbook describes major concepts, methods, and applications in computational linguistics in a way that undergraduates and non-specialists can comprehend. As described on Amazon, it&#8217;s a state-of-the-art reference to one of the most active and productive fields in linguistics. A wide range of linguists and researchers in fields such as informatics, artificial intelligence, language engineering, and cognitive science will find it interesting and practical. It begins with linguistic fundamentals, followed by an overview of current tasks, techniques, and tools in Natural Language Processing that target more experienced computational language researchers. Whether you&#8217;re a non-specialist or post-doctoral worker, this book will be useful.</p></li><li><p>&#8220;Foundations of Statistical Natural Language Processing&#8221; by Christopher Manning and Hinrich Schuetze.Another book that hails from Stanford educators, this one is written by Jurafsky&#8217;s colleague, Christopher Manning. They&#8217;ve taught the popular NLP introductory course at Stanford. Manning&#8217;s co-author is a professor of Computational Linguistics at the German Ludwig-Maximilians-Universit&#228;t. The book provides an introduction to statistical methods for NLP and a decent foundation to comprehend new NLP methods and support the creation of NLP tools. Mathematical and linguistic foundations, plus statistical methods, are equally represented in a way that supports readers in creating language processing applications.</p></li><li><p>&#8220;Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit&#8221; by Steven Bird, Ewan Klein, and Edward Loper. This book is a helpful introduction to the NLP field with a focus on programming. If you want to have a practical source on your shelf or desk, whether you&#8217;re a NLP beginner, computational linguist or AI developer, it contains hundreds of fully-worked examples and graded exercises that bring NLP to life. It can be used for individual study, as a course textbook when studying NLP or computational linguistics, or in complement with artificial intelligence, text mining, or corpus linguistics courses. Curious about the Python programming language? It will walk you through creating Python programs that parse unstructured data like language and recommends downloading Python and the Natural Language Toolkit.&nbsp;</p></li><li><p>&#8220;Big Data Analytics Methods: Modern Analytics Techniques for the 21st Century: The Data Scientist&#8217;s Manual to Data Mining, Deep Learning &amp; Natural Language Processing&#8221; by Peter Ghavami. This book might seem daunting to a NLP newcomer, but it&#8217;s useful as a comprehensive manual for those familiar with NLP and how big data relates in today&#8217;s world. It also works as a helpful reference for data scientists, analysts, business managers, and Business Intelligence practitioners. With more than a hundred analytics techniques and methods included, we think this will be a favorite for seasoned analytics practitioners. Chapters cover everything from machine learning to predictive modeling and cluster analysis. Data science topics including data visualization, prediction, and regression analysis, plus NLP-related fields such as neural networks, deep learning, and artificial intelligence are also discussed. These come with a broad explanation, but Peter goes into more detail about terminology and mathematical foundations, too.</p></li><li><p>"GPT-3: Building Innovative NLP Products Using Large Language Models" by Ben Awad. This book provides a hands-on approach to building innovative NLP products using GPT-3, one of the largest and most powerful language models available today.&nbsp;</p></li><li><p>Hands-On Generative AI with Transformers and Diffusion Models by Daniel Shiffman. This book offers a practical guide to working with generative AI, including the use of transformers and diffusion models. It's ideal for those who prefer learning by doing.&nbsp;</p></li><li><p>Representation Learning for Natural Language Processing by Zhiyuan Liu, Yankai Lin, and Maosong Sun. Published in 2021, this book offers an overview of recent advances in representation learning theory, algorithms, and applications for NLP. It covers everything from word embeddings to pre-trained language models, and is suitable for advanced undergraduate and graduate students, post-doctoral fellows, researchers, lecturers, and industrial engineers.</p></li></ol><div><hr></div><h3>Online courses</h3><h4>Natural Language Processing (NLP)</h4><ol><li><p>"Natural Language Processing" by DeepLearning.AI on Coursera.</p></li><li><p>&nbsp;"Deep Learning for Natural Language Processing" from deeplearning.ai.</p></li><li><p>"Generative AI with Large Language Models" also by DeepLearning.AI on Coursera.</p></li><li><p>"Natural Language Processing with Classification and Vector Spaces" by DeepLearning.AI on Coursera.</p></li><li><p>"Natural Language Processing with Sequence Models" by DeepLearning.AI on Coursera.</p></li><li><p>"Natural Language Processing with Attention Models" by DeepLearning.AI on Coursera.</p></li><li><p>"Python for Everybody" by University of Michigan on Coursera.</p></li><li><p>"Natural Language Processing on Google Cloud" by Google Cloud on Coursera.</p></li><li><p>"Foundations of Natural Language Processing" by Stanford University on edX.</p></li><li><p>"Natural Language Processing" by University of California, Irvine on Udemy.</p></li><li><p>"Natural Language Processing with Python" by University of Helsinki on FutureLearn.</p></li></ol><h4>Large Language Models (LLM)</h4><ol><li><p>"Introduction to Large Language Models" by DeepLearning.AI on Coursera.</p></li><li><p>"Finetuning Large Language Models" by DeepLearning.AI on Coursera.</p></li><li><p>"Prompt Engineering for ChatGPT" by Vanderbilt University on Coursera.</p></li><li><p>"ChatGPT Prompt Engineering for Developers" by DeepLearning.AI on Coursera.</p></li><li><p>"Building Systems with the ChatGPT API" by DeepLearning.AI on Coursera.</p></li><li><p>"LangChain Chat with Your Data" by DeepLearning.AI on Coursera.</p></li><li><p>"ChatGPT Advanced Data Analysis" by Vanderbilt University on Coursera.</p></li><li><p>"Understanding Large Language Models" by Google Cloud on YouTube.</p></li><li><p>"Large Language Models" by DeepMind on YouTube.</p></li><li><p>"Understanding Large Language Models" by Stanford University on YouTube.</p></li></ol><div><hr></div><p><strong>NLP Online Courses:</strong></p><p>Find the link below to access the above courses:</p><ul><li><p><a href="https://www.coursera.org/courses?query=nlp">https://www.coursera.org/courses?query=nlp</a></p></li><li><p>https://www.coursera.org/courses?query=large%20language%20models</p></li></ul><div><hr></div><h3>Conferences</h3><ol><li><p>Association for Computational Linguistics (ACL): This is one of the largest international conferences in computational linguistics, covering various areas including NLP and LLM.</p></li></ol><ol start="2"><li><p>Conference on Neural Information Processing Systems (NeurIPS): While not exclusively focused on NLP, many NLP and LLM researchers present their work at this conference.</p></li><li><p>Conference on Computational Natural Language Learning (CoNLL): CoNLL is a yearly conference organized by SIGNLL (ACL's Special Interest Group on Natural Language Learning). The focus of CoNLL is on theoretically, cognitively, and scientifically motivated approaches to computational linguistics. This includes computational learning theory and other techniques for theoretical analysis of machine learning models for NLP, models of language evolution and change, computational simulation and analysis of findings from psycholinguistic and neurolinguistic experiments, and more.&nbsp;</p></li></ol><ol start="3"><li><p>International Conference on Learning Representations (ICLR): This conference is a leading venue for machine learning research, and thus includes presentations on NLP and LLM.</p></li></ol><ol start="4"><li><p>Empirical Methods in Natural Language Processing (EMNLP): EMNLP is a major conference for empirical research in NLP, and it often features presentations on LLM.</p></li></ol><ol start="5"><li><p>NAACL HLT (Human Language Technology): This conference focuses specifically on NLP, and it's a good place to find up-to-date research in this field.</p></li></ol><ol start="6"><li><p>The Workshop on Deep Learning in Natural Language Processing (DeepNLP): As the name suggests, this workshop focuses on deep learning methods in NLP, which are often applied to LLM.</p></li></ol><ol start="7"><li><p>The Conference on Machine Learning and Knowledge Discovery in Databases (KDD): KDD includes sessions on NLP and LLM, making it another valuable resource.</p></li><li><p>AI In Finance Summit: This conference focuses on insights and technical use cases from AI specialists and data scientists in Financial Services. It covers topics such as the current AI landscape in finance, trends, the ROI of AI in financial services, AI advancements in fintech, and niche ML applications.</p></li><li><p>Arize:Observe: This is a one-day virtual summit dedicated to LLM observability. The conference includes presentations and panels from thought leaders and AI teams across industries. Topics covered range from performance monitoring and troubleshooting, data quality and troubleshooting, AI observability, explainability, ROI, and more.</p></li><li><p>NLP Summit: Organized by John Snow Labs, this summit is a global gathering of AI practitioners and researchers focusing on NLP technologies. The summit covers the latest developments in the field, including the application of LLMs.</p></li></ol><div><hr></div><h3>NLP Datasets</h3><ol><li><p>Google's Books Corpus: A massive dataset of text from books and articles.</p></li><li><p>Common Crawl: A collection of text and code crawled from the internet.</p></li><li><p>Stanford Natural Language Inference (SNLI) Corpus: A dataset for evaluating natural language inference.</p></li><li><p>WordNet: A lexical database of English words and their relationships.</p></li><li><p>IMDB Reviews: This dataset consists of over 25,000 movie reviews from the IMDB website. It's perfect for binary sentiment classification tasks, where the goal is to classify a review as positive or negative.</p></li><li><p>Multi-Domain Sentiment Analysis Dataset: This dataset includes a vast array of Amazon product reviews. It's particularly useful for sentiment analysis tasks, where the aim is to determine whether a piece of text expresses a positive or negative sentiment.</p></li><li><p>Stanford Sentiment Treebank: This dataset contains over 10,000 movie reviews from Rotten Tomatoes. It's excellent for sentiment analysis tasks, especially when dealing with longer phrases.</p></li><li><p>Sentiment140: This dataset comprises over 160,000 tweets, formatted within six fields including tweet data, query, text, polarity, ID, and user. It's suitable for sentiment analysis tasks, particularly for analyzing sentiments expressed in social media posts.</p></li><li><p>20 Newsgroups: This dataset includes 20,000 documents that cover 20 newsgroups and subjects. It's ideal for text classification tasks, where the goal is to categorize a piece of text into one of several predefined classes.</p></li><li><p>Reuters News Dataset: This dataset, which appeared in 1987, has been labeled, indexed, and compiled for use in machine learning. It's excellent for text classification tasks, particularly for categorizing news articles into different topics.</p></li><li><p>ArXiv: This massive 270 GB dataset features all arXiv research papers in full text. It's suitable for text classification tasks, particularly for categorizing scientific papers into different research areas.</p></li><li><p>The WikiQA Corpus: This publicly-available Q&amp;A dataset was initially compiled to aid in all open-domain question answering research. It's ideal for question answering tasks, where the goal is to generate a response given a question.</p></li><li><p>UCI&#8217;s Spambase: This dataset was created by a team at HP (Hewlett-Packard) to help create a spam filter. It contains a litany of emails previously labeled as spam by users. It's perfect for spam detection tasks, where the goal is to identify whether a piece of text is spam or not.</p></li><li><p>Yelp Reviews: This Yelp dataset features 8.5M+ reviews of over 160,000 businesses. It also has 200,000+ pictures and spans across 8 major metropolitan areas. It's suitable for sentiment analysis tasks, particularly for analyzing sentiments expressed in customer reviews.</p></li></ol><p>Here are some of the latest datasets used for training Large Language Models (LLMs):</p><p>1. RefinedWeb: This is a massive corpus of deduplicated and filtered tokens from the Common Crawl dataset. With more than 5 trillion tokens of textual data, of which 600 billion are made publicly available, it was developed as an initiative to train the Falcon-40B model with smaller-sized but high-quality datasets.</p><p>2. The Pile: This is an 800 GB corpus that enhances a model&#8217;s generalization capability across a broader context. It was curated from 22 diverse datasets, mostly from academic or professional sources. The Pile was instrumental in training various LLMs, including GPT-Neo, LLaMA, and OPT.</p><p>3. Starcoder Data: This is a programming-centric dataset built from 783 GB of code written in 86 programming languages. It also contains 250 billion tokens extracted from GitHub and Jupyter Notebooks. Salesforce CodeGen, Starcoder, and StableCode were trained with Starcoder Data to enable better program synthesis.</p><p>4. BookCorpus: This dataset turned scraped data of 11,000 unpublished books into a 985 million-word dataset. It was initially created to align storylines in books to their movie interpretations. The dataset was used for training LLMs like RoBERTa, XLNET, and T5.</p><p>5. ROOTS: This is a 1.6TB multilingual dataset curated from text sourced in 59 languages. Created to train the BigScience Large Open-science Open-access Multilingual (BLOOM) language model. ROOTS uses heavily deduplicated and filtered data from Common Crawl, GitHub Code, and other crowdsourced initiatives..</p><p>6. Wikipedia: The Wikipedia dataset is curated from cleaned text data derived from the Wikipedia site and presented in all languages. The default English Wikipedia dataset contains 19.88 GB of vast examples of complete articles that help with language modeling tasks. It was used to train larger models like Roberta, XLNet, and LLaMA.</p><p>7.Databricks-dolly-15k: This is a dataset for LLM finetuning that features &gt;15,000 instruction-pairs written by thousands of DataBricks employees. It is similar to those used to train systems like InstructGPT and ChatGPT.</p><p>8. OpenAssistant Conversations: This is another dataset for fine tuning pretraining LLMs on a collection of ChatGPT assistant-like dialogues that have been created and annotated by humans, encompassing 161,443 messages in 35 diverse languages, along with 461,292 quality assessments. These are organized in more than 10,000 thoroughly annotated conversation trees.</p><p>9. RedPajama: This is an open-source dataset for pretraining LLMs similar to LLaMA Meta's state-of-the-art LLaMA model. The goal of this project is to create a capable open-source competitor to the most popular LLMs, which are currently closed commercial models or only partially open [Source 4](https://magazine.sebastianraschka.com/p/ahead-of-ai-8-the-latest-open-source).</p><p>10. Pythia: This is an 800GB dataset of diverse texts for 300 B tokens (~1 epoch on regular Pile, ~1.5 epochs on deduplicated Pile). The recent instruction-finetuned open-source Dolly v2 LLM model used Pythia as the base (foundation) model.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HKHJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HKHJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png 424w, https://substackcdn.com/image/fetch/$s_!HKHJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png 848w, https://substackcdn.com/image/fetch/$s_!HKHJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png 1272w, https://substackcdn.com/image/fetch/$s_!HKHJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HKHJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png" width="1388" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1388,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:48696,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!HKHJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png 424w, https://substackcdn.com/image/fetch/$s_!HKHJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png 848w, https://substackcdn.com/image/fetch/$s_!HKHJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png 1272w, https://substackcdn.com/image/fetch/$s_!HKHJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0d0963c-8792-4ab3-9a92-600a341a6f2c_1388x344.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><div><hr></div><h3>NLP Libraries</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!flfM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!flfM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png 424w, https://substackcdn.com/image/fetch/$s_!flfM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png 848w, https://substackcdn.com/image/fetch/$s_!flfM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png 1272w, https://substackcdn.com/image/fetch/$s_!flfM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!flfM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png" width="872" height="588" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:588,&quot;width&quot;:872,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:138327,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!flfM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png 424w, https://substackcdn.com/image/fetch/$s_!flfM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png 848w, https://substackcdn.com/image/fetch/$s_!flfM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png 1272w, https://substackcdn.com/image/fetch/$s_!flfM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe22cf48d-b846-4cc7-944f-749ac97aa635_872x588.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ol><li><p>AllenNLP: This is an Apache 2.0 NLP research library, built on PyTorch, for developing state-of-the-art deep learning models on a wide variety of linguistic tasks. It provides a broad collection of existing model implementations that are well documented and engineered to a high standard, making them a great foundation for further research. AllenNLP offers a high-level configuration language to implement many common approaches in NLP, such as transformer experiments, multi-task training, vision+language tasks, fairness, and interpretability. This allows experimentation on a broad range of tasks purely through configuration, so you can focus on the important questions in your research.</p></li><li><p>NLTK: NLTK &#8212; the Natural Language Toolkit &#8212; is a suite of open-source Python modules, data sets, and tutorials supporting research and development in Natural Language Processing. It provides easy-to-use interfaces to over 50 corpora and lexical resources such as WordNet, along with a suite of text processing libraries for classification, tokenization, stemming, tagging, parsing, and semantic reasoning, wrappers for industrial-strength NLP libraries.</p></li><li><p>NLP Architect: NLP Architect is an open-source Python library for exploring state-of-the-art deep learning topologies and techniques for optimizing Natural Language Processing and Natural Language Understanding Neural Networks. It&#8217;s a library designed to be flexible, easy to extend, allowing for easy and rapid integration of NLP models in applications, and to showcase optimized models.</p></li><li><p>PyTorch-NLP: PyTorch-NLP is a library of basic utilities for PyTorch NLP. It extends PyTorch to provide you with basic text data processing functions.</p></li><li><p>John Snow Labs: John Snow Labs' NLP &amp; LLM ecosystem include software libraries for state-of-the-art AI at scale, Responsible AI, No-Code AI, and access to over 20,000 models for Healthcare, Legal, Finance, and Visual NLP.</p></li><li><p>PyNLPl: Pronounced as &#8216;pineapple,&#8217; PyNLPl is a Python library for Natural Language Processing. It contains a collection of custom-made Python modules for Natural Language Processing tasks. One of the most notable features of PyNLPl is that it features an extensive library for working with FoLiA XML (Format for Linguistic Annotation). PyNLPl is segregated into different modules and packages, each useful for both standard and advanced NLP tasks. While you can use PyNLPl for basic NLP tasks like extraction of n-grams and frequency lists, and to build a simple language model, it also has more complex data types and algorithms for advanced NLP tasks.</p></li><li><p>Stanford CoreNLP: Stanford CoreNLP is a Java library for NLP processing. It provides a set of natural language analysis tools that include token and sentence boundaries, parts of speech, named entities, numeric and time values, dependency and constituency parses, coreference, sentiment, quote attributions, and relations. It's a comprehensive library for NLP tasks and is widely used in academia and industry.</p></li><li><p>TensorFlow: TensorFlow is an open-source library developed by Google Brain Team. It's used for numerical computation, particularly well suited for large-scale Machine Learning. It provides a flexible platform for defining and running machine learning algorithms. It supports a wide range of tasks including computer vision, natural language processing, and more.</p></li><li><p>Hugging Face Transformers: Hugging Face Transformers is a state-of-the-art Natural Language Processing library for TensorFlow 2.0 and PyTorch. It provides thousands of pre-trained models to perform tasks on texts such as classification, information extraction, generation, etc. It's widely used in the field of NLP due to its ease of use and the availability of pre-trained models.</p></li><li><p>Keras: Keras is a high-level neural networks API, written in Python and capable of running on top of TensorFlow, CNTK, or Theano. It allows for easy and fast prototyping and supports both convolutional networks and recurrent networks, as well as combinations of the two.</p></li><li><p>Pattern: Pattern is a web mining module for Python. It includes tools for natural language processing, machine learning, and graph theory. It's used for tasks such as part-of-speech tagging, named entity recognition, sentiment analysis, text classification, and topic modeling.</p></li><li><p>Polyglot: Polyglot is a Python library that implements various algorithms for natural language processing. It supports a wide range of languages and tasks, including morphological analysis, syntax parsing, machine translation, and more.</p></li><li><p>gensim: gensim is a Python library for topic modelling, document indexing, and similarity retrieval with large corpora. It's used for creating vector space models, performing topic modelling, and building similarity retrieval systems.</p></li></ol><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1UeB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1UeB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png 424w, https://substackcdn.com/image/fetch/$s_!1UeB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png 848w, https://substackcdn.com/image/fetch/$s_!1UeB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png 1272w, https://substackcdn.com/image/fetch/$s_!1UeB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1UeB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png" width="1410" height="1136" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1136,&quot;width&quot;:1410,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:174091,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!1UeB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png 424w, https://substackcdn.com/image/fetch/$s_!1UeB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png 848w, https://substackcdn.com/image/fetch/$s_!1UeB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png 1272w, https://substackcdn.com/image/fetch/$s_!1UeB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b86889-2c61-4036-8e05-fcb1a7c9f7bf_1410x1136.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><ol><li></li></ol><div><hr></div><h2>Summary</h2><p>This article introduces Natural Language Processing (NLP), a pivotal field in computer science that bridges the gap between computers and human language, enabling machines to understand, generate, translate, and even write human-quality text. NLP has seen remarkable advancements, making it possible for computers to comprehend programming languages, biological sequences, and even the nuances of human language. The article delves into the fundamental concepts of NLP, such as tokenization, stemming, and tagging, and explores its applications in tasks like machine translation, sentiment analysis, and text summarization.It also offers resources for further exploration, including popular NLP datasets, libraries, and online courses. By the end of this article, now you will have a grasp of the components of an NLP system and the complexities involved in developing these systems.</p><div><hr></div><h2>Quiz Question</h2><ol><li><p>What is the primary purpose of Natural Language Processing (NLP)?</p></li></ol><blockquote><p>&nbsp;A) To process binary data</p><p>B) To enhance user experience on websites</p><p>C) To enable computers to understand and generate human language</p><p>D) To optimize database performance</p></blockquote><ol start="2"><li><p>Which of the following is NOT a common task in NLP?</p></li></ol><blockquote><p>A) Sentiment analysis</p><p>B) Machine translation</p><p>C) Speech recognition</p><p>D) Image captioning</p></blockquote><ol start="3"><li><p>What is the process of breaking down text into smaller pieces called?</p></li></ol><blockquote><p>A) Parsing</p><p>B) Tagging</p><p>C) Tokenization</p><p>D) Stemming</p></blockquote><ol start="4"><li><p>What challenge does NLP face due to the ambiguity of human language?</p></li></ol><blockquote><p>A) Inability to understand context</p><p>B) Inability to handle synonyms</p><p>C) Inability to understand slang</p><p>D) All of the above</p></blockquote><p><strong>Correct Answers:</strong></p><ol><li><p>C</p></li><li><p>D</p></li><li><p>C</p></li><li><p>D</p></li></ol><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://aboniasojasingarayar.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://aboniasojasingarayar.substack.com/subscribe?"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[Understanding Kubernetes: Comprehensive Practical Guide]]></title><description><![CDATA[Components, Importance and Advantages]]></description><link>https://aboniasojasingarayar.substack.com/p/understanding-kubernetes-comprehensive-practical-guide-461bf0372330</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/understanding-kubernetes-comprehensive-practical-guide-461bf0372330</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Thu, 20 Jun 2024 07:32:33 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/942ddb9b-5e6b-493d-97e1-9c4b5a27799f_800x566.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Components, Importance and Advantages</h4><p>K<strong>ubernetes</strong>, often abbreviated as K8s, has emerged as a cornerstone technology in modern cloud-native application development and deployment. This article delves into the essence of Kubernetes, highlighting its necessity, significance, benefits, and key components, tailored for DevOps and MLOps engineers.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!egZS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!egZS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png 424w, https://substackcdn.com/image/fetch/$s_!egZS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png 848w, https://substackcdn.com/image/fetch/$s_!egZS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png 1272w, https://substackcdn.com/image/fetch/$s_!egZS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!egZS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!egZS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png 424w, https://substackcdn.com/image/fetch/$s_!egZS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png 848w, https://substackcdn.com/image/fetch/$s_!egZS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png 1272w, https://substackcdn.com/image/fetch/$s_!egZS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F155755aa-2f9b-4d98-95d7-78370f6a32a1_800x566.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Courtesy&#8202;&#8212;&#8202;Wikipedia</figcaption></figure></div><h3>Introduction to Kubernetes</h3><p>Kubernetes is an open-source platform designed to automate deploying, scaling, and operating application containers across clusters of hosts. It groups containers that make up an application into logical units for easy management and discovery. Kubernetes was originally developed by Google and is now maintained by the Cloud Native Computing Foundation (CNCF).</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oBPm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oBPm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg 424w, https://substackcdn.com/image/fetch/$s_!oBPm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg 848w, https://substackcdn.com/image/fetch/$s_!oBPm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!oBPm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oBPm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oBPm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg 424w, https://substackcdn.com/image/fetch/$s_!oBPm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg 848w, https://substackcdn.com/image/fetch/$s_!oBPm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!oBPm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1850705-d2ca-47d9-b46a-137140489d15_800x341.jpeg 1456w" sizes="100vw"></picture><div></div></div></a></figure></div><blockquote><p>Documentation: <a href="https://kubernetes.io/docs/home/">https://kubernetes.io/docs/home/</a></p></blockquote><blockquote><p>Github: <a href="https://github.com/kubernetes/kubernetes">https://github.com/kubernetes/kubernetes</a></p></blockquote><h3>Understanding MLOps and&nbsp;DevOps</h3><h4>MLOps: Bridging the Gap Between Data Science and Production</h4><p>MLOps, or Machine Learning Operations, is an extension of DevOps principles applied to machine learning (ML) workflows. It aims to streamline the process of moving ML models from development to production, ensuring scalability, maintainability, and delivering expected business value. MLOps emphasizes collaboration between data scientists, operations teams, and other stakeholders, leveraging tools and practices that facilitate the management and monitoring of ML models throughout their lifecycle.</p><h4>DevOps: Enhancing Software Development and&nbsp;Delivery</h4><p>DevOps, a combination of software development (Dev) and IT operations (Ops), focuses on optimizing the development, deployment, and maintenance of software applications. It promotes continuous integration and continuous delivery (CI/CD), fostering a culture of collaboration and communication between development and operations teams. The goal is to accelerate software delivery, increase quality, and enhance responsiveness to market demands.</p><h3><strong>Why Kubernetes?</strong></h3><p>The digital landscape is increasingly characterized by microservices architecture, where applications are broken down into smaller, independent services. Managing these services efficiently requires orchestration, which is where Kubernetes shines. Its ability to manage containerized applications across a cluster of machines offers several compelling reasons:</p><p>- Scalability: Easily scale applications based on demand.</p><p>- Self-healing: Automatically restart failed containers, replace containers, kill containers that don&#8217;t respond to health checks, and adjust the number of containers in a replica set.</p><p>- Load Balancing: Efficiently distribute network traffic across multiple pods.</p><p>- Automated Rollouts and Rollbacks: We can record the desired state of our deployed containers using Kubernetes, and it can change the actual state to match the desired state at a large scale and speed.</p><p>- Storage Orchestration: Abstracting storage into a portable resource just like compute nodes.</p><h3>Importance of Kubernetes</h3><p>In the realm of DevOps and MLOps, Kubernetes stands out for several reasons:</p><p>- DevOps Efficiency: Kubernetes simplifies the process of managing complex deployments, enabling teams to focus on developing features rather than managing infrastructure.</p><p>- MLOps Integration: With its support for custom metrics and autoscaling, Kubernetes facilitates the deployment and scaling of machine learning models, making it a crucial component in MLOps pipelines.</p><p>-Cloud-Native Applications: Kubernetes is inherently designed for cloud environments, supporting both public and private clouds, making it ideal for building and running cloud-native applications.</p><h3>Key Components of Kubernetes</h3><p>Kubernetes is composed of several core components, each playing a crucial role in managing applications and infrastructure.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OK5D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OK5D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OK5D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OK5D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OK5D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OK5D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OK5D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OK5D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OK5D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OK5D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2f60cc4-e530-4df4-88fd-2fe59250153b_800x378.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Image Credit: <a href="https://www.linux.com/news/learn/chapter/intro-to-kubernetes/2017/4/what-makes-kubernetes-cluster">Linux.com</a></figcaption></figure></div><p>- Pods: The smallest deployable units that can be created and managed in Kubernetes. A Pod represents a single instance of a running process in a cluster and can contain one or more containers.</p><p>- Services: An abstraction which defines a logical set of Pods and a policy by which to access them, sometimes called a micro-service.</p><p>- Deployments: Manage the creation and scaling of replicated applications.</p><p>- DaemonSets: Ensure that all (or some) Nodes run a copy of a pod.</p><p>- StatefulSets: Manage the deployment and scaling of a set of Pods, providing guarantees about the ordering and uniqueness of these Pods.</p><p>- ConfigMaps and Secrets: Used to store configuration settings and sensitive information such as passwords, OAuth tokens, and ssh keys.</p><p><strong>-</strong>Namespaces: Provide a scope for names and are intended for use in environments with many users spread across multiple teams, projects, or organizations.</p><h3>1. Master&nbsp;Nodes</h3><p>The master nodes make global decisions about the cluster, such as scheduling workloads onto nodes. They are responsible for maintaining the desired state of the cluster, such as which applications are running and where they should be deployed.</p><p><strong>Control Plane Components:</strong></p><p>- <strong>API Server</strong>: The front-end for the Kubernetes control plane. Processes REST operations.<br>&#8202;&#8212;&#8202;<strong>Controller Manager</strong>: Runs controllers, background processes that handle the overall state of the cluster.<br>&#8202;&#8212;&#8202;<strong>Scheduler</strong>: Assigns newly created pods to nodes.<br>&#8202;&#8212;&#8202;<strong>etcd</strong>: Stores all data needed to manage the cluster and configures the cluster&#8217;s state.</p><h3>2. Worker&nbsp;Nodes</h3><p>Worker nodes host the applications and provide resources such as compute, memory, storage, and networking. They receive instructions from the master nodes on what to run.</p><h4>Node Components:</h4><p>- <strong>Kubelet</strong>: An agent that runs on each node in the cluster. Ensures that containers are running in a pod.<br>&#8202;&#8212;&#8202;<strong>Container Runtime</strong>: Responsible for running containers. Examples include Docker, containerd, CRI-O.<br>&#8202;&#8212;<strong>&#8202;Kube-proxy</strong>: A network proxy that runs on each node in the cluster, maintaining network rules on nodes.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TxbK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TxbK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TxbK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TxbK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TxbK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TxbK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TxbK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg 424w, https://substackcdn.com/image/fetch/$s_!TxbK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg 848w, https://substackcdn.com/image/fetch/$s_!TxbK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!TxbK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e3effd1-b696-4ac5-9c35-a796569f66dc_800x534.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><h4>Deployments</h4><p>Deployments manage the deployment and scaling of applications. They describe the desired state of your application, including the number of replicas and the template for creating new pods.</p><h4>Volumes</h4><p>Volumes provide persistent storage for applications. Unlike ephemeral containers, volumes remain even after a container is terminated, allowing data to persist across container restarts.</p><h4>ConfigMaps and&nbsp;Secrets</h4><p>ConfigMaps and Secrets allow you to separate configuration artifacts and sensitive information (e.g., passwords, OAuth tokens) from image content to keep containerized applications portable.</p><h4>Ingress and Control&nbsp;IPs</h4><p>Ingress manages external access to the services in a cluster, providing HTTP and HTTPS routes. Control IPs, part of the networking model, play a crucial role in routing external traffic to the appropriate services within the cluster.</p><h4>Networking in Kubernetes</h4><p>Networking is a critical component of Kubernetes, enabling communication between pods and services within the cluster and with external networks. Kubernetes supports various network models, including overlay networks and third-party solutions.</p><p><strong>Service Discovery: </strong>Service discovery enables pods to find and communicate with each other using DNS names. Kubernetes provides a DNS service for this purpose.</p><p><strong>Load Balancing&nbsp;:</strong>Kubernetes uses a load balancer to distribute incoming traffic among the available pods. This ensures high availability and fault tolerance.</p><h4>Security in Kubernetes</h4><p>Security in Kubernetes is a broad topic covering authentication, authorization, network policies, and secure communication between pods and services.</p><h4>Role-Based Access Control&nbsp;(RBAC)</h4><p>RBAC allows fine-grained access management to Kubernetes resources. Users can be assigned roles with specific permissions.</p><h4>Network Policies</h4><p>Network policies define how groups of pods are allowed to communicate with various network endpoints and with other pods.</p><h4>Storage in Kubernetes</h4><p>Kubernetes supports both local and networked storage options. Persistent Volumes (PVs) represent physical storage resources in the cluster, while Persistent Volume Claims (PVCs) request those resources for use by pods.</p><h3>Minikube&#8202;&#8212;&#8202;single-node Kubernetes cluster on local&nbsp;machine</h3><p>To run Kubernetes locally, we can use <strong><a href="https://github.com/kubernetes/minikube">Minikube</a></strong>, a tool that allows we to create a single-node Kubernetes cluster on our personal computer. This setup is ideal for learning Kubernetes, developing applications, and testing Kubernetes configurations without affecting a production environment. Here&#8217;s how to get started with Minikube and perform some basic operations.</p><p>Step 1: Install Minikube</p><p>Before installing <strong><a href="https://github.com/kubernetes/minikube">Minikube</a></strong>, ensure we have <strong><a href="https://www.docker.com/">Docker</a></strong> or another container runtime installed on our system. The installation process varies depending on our operating system.</p><p>Step 2: Start Minikube</p><p>After installing Minikube, we can start a local Kubernetes cluster by running the following command in our terminal:</p><pre><code>minikube start</code></pre><p>This command initializes a single-node Kubernetes cluster. Minikube uses <strong><a href="https://www.virtualbox.org/">VirtualBox</a></strong> by default as the VM provider, but it can also work with other providers like Hyper-V on Windows or VMware Fusion on macOS.</p><p>Step 3: Verify the Cluster</p><p>Once Minikube starts, we can verify that our local Kubernetes cluster is running correctly by listing the nodes in the cluster:</p><pre><code>kubectl get nodes</code></pre><p>we should see output indicating that our node (`minikube`) is in the `Ready` state.</p><p>Step 4: Deploy an Application</p><p>With our local Kubernetes cluster running, we can deploy applications just as we would in a production environment. For example, to deploy a simple nginx web server, we can use the following command:</p><pre><code>kubectl run nginx - image=nginx - port=80</code></pre><p>This command creates a deployment named `nginx` using the `nginx` image and exposes port 80.</p><p>Step 5: Access the Application</p><p>To access the application running in our local Kubernetes cluster, we can use Minikube&#8217;s built-in tunnel feature:</p><pre><code>minikube tunnel</code></pre><p>Then, open a web browser and navigate to `http://localhost`. we should see the nginx welcome page served by our local Kubernetecluster.</p><p>- Stop Minikube: To stop the local Kubernetes cluster, use t command:</p><pre><code>minikube stop</code></pre><p>- Start Minikube with a Specific Version: If we need to test our application against a specific Kubernetes version, we can start Minikube with a custom version:</p><pre><code>minikube start - kubernetes-vsion=v1.22.0</code></pre><p>- Access the Kubernetes Dashboard: Minikube includes the Kubernetes Dashboard, a web-based UI for Kubernetes clusters. We can access it by running:</p><pre><code>minikube dashboard</code></pre><p>By following these steps, we can easily run Kubernetes locally on our machine using Minikube. This setup is invaluable for development, testing, and learning Kubernetes without the need for a full-scale cloud environment.</p><p>To complement the theoretical aspects of Kubernetes covered earlier, let&#8217;s delve into practical implementations using `<strong><a href="https://kubernetes.io/docs/tasks/tools/">kubectl</a></strong>`, the command-line interface for interacting with Kubernetes clusters. These commands are essential for managing various Kubernetes resources, ranging from pods and deployments to services and namespaces.</p><h3>Getting Started with Kubernetes</h3><p>First, ensure we have `<strong><a href="https://kubernetes.io/docs/tasks/tools/">kubectl</a></strong>` installed and configured to communicate with our Kubernetes cluster. This setup typically involves pointing `kubectl` to the correct cluster endpoint and authenticating with the cluster&#8217;s API server.</p><p>Checking Cluster Health</p><p>- <strong>Check Worker Nodes Status</strong></p><pre><code>kubectl get nodes</code></pre><p>Use `-o wide` for more detailed information about each node.</p><pre><code>kubectl get nodes -o wide</code></pre><p>Creating Resources</p><p>- <strong>Create a Pod</strong></p><pre><code>kubectl run my-first-pod - image stacksimplify/kubenginx:1.0.0</code></pre><p>This command creates a pod named `my-first-pod` using the specified image.</p><p>Listing and Describing Pods</p><p>- <strong>List Pods</strong></p><pre><code>kubectl get pods</code></pre><p>- <strong>Describe a Specific Pod</strong></p><pre><code>kubectl describe pod my-first-pod</code></pre><p>Interacting with Deployments and Services</p><p>- <strong>Dump Logs for a Deployment</strong></p><pre><code>kubectl logs deploy/my-deployment</code></pre><p>For multi-container deployments, specify the container name:</p><pre><code>kubectl logs deploy/my-deployment -c my-container</code></pre><p>- <strong>Port Forwarding</strong></p><p>Forward a local port to a service or deployment:</p><pre><code>kubectl port-forward svc/my-service 5000</code></pre><p>Or to a specific port within a deployment:</p><pre><code>kubectl port-forward deploy/my-deployment 5000:6000</code></pre><p>- <strong>Executing Commands in a Container</strong></p><p>Run a command inside a container of a pod:</p><pre><code>kubectl exec deploy/my-deployment - ls</code></pre><p>Managing Resources with YAML Files</p><p>- <strong>Apply a YAML File</strong></p><p>Apply the configuration defined in a YAML file:</p><pre><code>kubectl apply -f &lt;filename.yaml&gt;</code></pre><p>- <strong>Delete Resources Defined in a YAML File</strong></p><p>Delete the resources specified in a YAML file:</p><pre><code>kubectl delete -f &lt;filename.yaml&gt;</code></pre><p>- <strong>Port Forwarding to a Pod</strong></p><p>Forward a local port to a specified port in a pod:</p><pre><code>kubectl port-forward &lt;pod name&gt; &lt;local port&gt;:&lt;remote port&gt;</code></pre><p>- <strong>Listing Logs for a Pod</strong></p><p>View logs for a pod, specifying the container name for multi-container pods:</p><pre><code>kubectl logs &lt;pod name&gt;</code></pre><p>To learn more with kubernetes Cheatsheet by <strong><a href="https://www.scaleway.com/en/docs/static/be9a6e5821a4e8e268c7c5bd3624e256/scaleway-kubernetes-cheatsheet.pdf">ScaleWay</a></strong>:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NPLR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NPLR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png 424w, https://substackcdn.com/image/fetch/$s_!NPLR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png 848w, https://substackcdn.com/image/fetch/$s_!NPLR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png 1272w, https://substackcdn.com/image/fetch/$s_!NPLR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NPLR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NPLR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png 424w, https://substackcdn.com/image/fetch/$s_!NPLR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png 848w, https://substackcdn.com/image/fetch/$s_!NPLR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png 1272w, https://substackcdn.com/image/fetch/$s_!NPLR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65d89201-e24b-4a19-8e00-8ead72d56d3d_800x488.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">K8s Cheatsheet</figcaption></figure></div><h3>Advantages of Using Kubernetes</h3><p>- Portability Across Environments: Containers packaged with Kubernetes can run anywhere, from laptops to massive clusters, providing consistency across different environments.</p><p>- Microservices Architecture Support: Kubernetes excels in managing microservices, offering tools for service discovery, load balancing, and fault tolerance.</p><p>- Community and Ecosystem: Being an open-source project, Kubernetes benefits from a vast community and ecosystem, including extensive documentation, plugins, and integrations with other tools.</p><h3>Conclusion</h3><p>Kubernetes has revolutionized how applications are deployed, scaled, and managed in the cloud. Its importance lies in its ability to handle the complexities of modern application architectures, offering a robust platform for DevOps and MLOps practices. By understanding its components and advantages, engineers can harness Kubernetes&#8217; power to build resilient, scalable, and efficient applications.</p><blockquote><p>Thanks for&nbsp;Reading</p></blockquote><blockquote><p><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></p></blockquote><blockquote><p><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></p></blockquote><blockquote><p><strong>Visit my technical channel on</strong> <strong><a href="http://www.youtube.com/@aboniasojasingarayar3097">Youtube</a></strong></p></blockquote>]]></content:encoded></item><item><title><![CDATA[vLLM and PagedAttention: A Comprehensive Overview]]></title><description><![CDATA[Easy, Fast, and Cheap LLM Serving]]></description><link>https://aboniasojasingarayar.substack.com/p/vllm-and-pagedattention-a-comprehensive-overview-20046d8d0c61</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/vllm-and-pagedattention-a-comprehensive-overview-20046d8d0c61</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Wed, 15 May 2024 07:32:02 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/681c6201-cc76-43ba-b528-86699e3fab4e_1200x590.gif" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Easy, Fast, and Cheap LLM&nbsp;Serving</h4><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XYP2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XYP2!,w_424,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif 424w, https://substackcdn.com/image/fetch/$s_!XYP2!,w_848,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif 848w, https://substackcdn.com/image/fetch/$s_!XYP2!,w_1272,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif 1272w, https://substackcdn.com/image/fetch/$s_!XYP2!,w_1456,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XYP2!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XYP2!,w_424,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif 424w, https://substackcdn.com/image/fetch/$s_!XYP2!,w_848,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif 848w, https://substackcdn.com/image/fetch/$s_!XYP2!,w_1272,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif 1272w, https://substackcdn.com/image/fetch/$s_!XYP2!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3bf6769-923a-4bac-ab28-c250f62f946e_1200x590.gif 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Source&#8202;&#8212;&#8202;vllm.ai</figcaption></figure></div><p>vLLM is a library designed to enhance the efficiency and performance of Large Language Model (LLM) inference and serving. Developed at UC Berkeley, vLLM introduces PagedAttention, a novel attention algorithm that significantly optimizes memory management for attention keys and values. This innovation not only boosts throughput but also enables continuous batching of incoming requests, fast model execution with CUDA/HIP graph, and supports various decoding algorithms including parallel sampling and beam search. vLLM is compatible with both NVIDIA and AMD GPUs, and it seamlessly integrates with popular Hugging Face models, making it a versatile tool for developers and researchers alike.</p><h3>PagedAttention: The Key Technique</h3><p>PagedAttention is the heart of vLLM&#8217;s performance enhancements. It addresses the critical issue of memory management in LLM serving by partitioning the KV cache into blocks, allowing for non-contiguous storage of keys and values in memory. This approach not only optimizes memory usage, reducing waste by up to 96%, but also enables efficient memory sharing, significantly reducing the memory overhead of complex sampling algorithms. PagedAttention&#8217;s memory management strategy is inspired by the concept of virtual memory and paging in operating systems, offering a flexible and efficient way to manage memory resources.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0LbV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0LbV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png 424w, https://substackcdn.com/image/fetch/$s_!0LbV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png 848w, https://substackcdn.com/image/fetch/$s_!0LbV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png 1272w, https://substackcdn.com/image/fetch/$s_!0LbV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0LbV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/994609f7-53f5-4406-9624-720e240835ca_800x437.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0LbV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png 424w, https://substackcdn.com/image/fetch/$s_!0LbV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png 848w, https://substackcdn.com/image/fetch/$s_!0LbV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png 1272w, https://substackcdn.com/image/fetch/$s_!0LbV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F994609f7-53f5-4406-9624-720e240835ca_800x437.png 1456w" sizes="100vw"></picture><div></div></div></a><figcaption class="image-caption">vLLM System overview&#8202;&#8212;&#8202;arxiv.2309.06180</figcaption></figure></div><h3>vLLM&#8217;s Features and Capabilities</h3><ul><li><p><strong>High Throughput and Memory Efficiency</strong>: vLLM delivers state-of-the-art serving throughput, making it an ideal choice for applications requiring high performance and low latency.</p></li><li><p><strong>Continuous Batching of Requests</strong>: vLLM efficiently manages incoming requests, allowing for continuous batching and processing.</p></li><li><p><strong>Fast Model Execution</strong>: Utilizing CUDA/HIP graph, vLLM ensures fast execution of models, enhancing the overall performance of LLM serving.</p></li><li><p><strong>Quantization Support</strong>: vLLM supports various quantization techniques, including GPTQ, AWQ, SqueezeLLM, and FP8 KV Cache, to further optimize model performance and reduce memory footprint.</p></li><li><p><strong>Optimized CUDA Kernels</strong>: vLLM includes optimized CUDA kernels for enhanced performance on NVIDIA GPUs.</p></li><li><p><strong>Tensor Parallelism Support</strong>: For distributed inference, vLLM offers tensor parallelism support, facilitating scalable and efficient model serving across multiple GPUs.</p></li><li><p><strong>Streaming Outputs</strong>: vLLM supports streaming outputs, allowing for real-time processing and delivery of model outputs.</p></li><li><p><strong>OpenAI-Compatible API Server</strong>: vLLM can be used to start an OpenAI API-compatible server, making it easy to integrate with existing systems and workflows.</p></li></ul><blockquote><p>Documenation and Paper</p></blockquote><p><strong><a href="https://github.com/vllm-project/vllm" title="https://github.com/vllm-project/vllm">GitHub - vllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for&#8230;</a></strong><a href="https://github.com/vllm-project/vllm" title="https://github.com/vllm-project/vllm"><br></a><em><a href="https://github.com/vllm-project/vllm" title="https://github.com/vllm-project/vllm">A high-throughput and memory-efficient inference and serving engine for LLMs - vllm-project/vllm</a></em><a href="https://github.com/vllm-project/vllm" title="https://github.com/vllm-project/vllm">github.com</a></p><p><strong><a href="https://vllm.ai/" title="https://vllm.ai/">vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention</a></strong><a href="https://vllm.ai/" title="https://vllm.ai/"><br></a><em><a href="https://vllm.ai/" title="https://vllm.ai/">GitHub | Documentation | Paper</a></em><a href="https://vllm.ai/" title="https://vllm.ai/">vllm.ai</a></p><h3>Integration with Hugging Face&nbsp;Models</h3><p>vLLM seamlessly supports a wide range of Hugging Face models, including Aquila, Baichuan, BLOOM, ChatGLM, DeciLM, Falcon, Gemma, GPT-2, GPT BigCode, GPT-J, GPT-NeoX, InternLM, Jais, LLaMA &amp; LLaMA-2, Mistral, MPT, OLMo, OPT, Orion, Phi, Qwen, Qwen2, StableLM, Starcoder2, and Yi. This broad compatibility ensures that vLLM can be used with a vast array of LLM architectures, making it a versatile tool for developers and researchers working with different model types.</p><pre><code># 1. vLLM - Offline Batch Inference
from vllm import LLM

# Sample prompts.
prompts = ["Hello, my name is", "Capital of France is"] 
# Create an LLM with HF. 
llm = LLM(model="gpt2") 
# Generate texts from the prompts. 
outputs = llm.generate(prompts)</code></pre><h3>Getting Started with&nbsp;vLLM</h3><p>To get started with vLLM, you can install it via pip and use it for both offline inference and online serving. For offline inference, you can import the <code>LLM</code> class from vLLM and generate texts from prompts. For online serving, vLLM can be used to start an OpenAI API-compatible server, allowing you to query the server in the same format as the OpenAI API. This ease of use, combined with vLLM's powerful features, makes it an attractive option for developers looking to leverage LLMs in their applications.</p><p><strong><a href="https://colab.research.google.com/gist/Abonia1/9a68f8e7e0772d8f60fbad0970885efb/vllm-inference-engine.ipynb" title="https://colab.research.google.com/gist/Abonia1/9a68f8e7e0772d8f60fbad0970885efb/vllm-inference-engine.ipynb">Google Colaboratory</a></strong><a href="https://colab.research.google.com/gist/Abonia1/9a68f8e7e0772d8f60fbad0970885efb/vllm-inference-engine.ipynb" title="https://colab.research.google.com/gist/Abonia1/9a68f8e7e0772d8f60fbad0970885efb/vllm-inference-engine.ipynb"><br></a><em><a href="https://colab.research.google.com/gist/Abonia1/9a68f8e7e0772d8f60fbad0970885efb/vllm-inference-engine.ipynb" title="https://colab.research.google.com/gist/Abonia1/9a68f8e7e0772d8f60fbad0970885efb/vllm-inference-engine.ipynb">Edit description</a></em><a href="https://colab.research.google.com/gist/Abonia1/9a68f8e7e0772d8f60fbad0970885efb/vllm-inference-engine.ipynb" title="https://colab.research.google.com/gist/Abonia1/9a68f8e7e0772d8f60fbad0970885efb/vllm-inference-engine.ipynb">colab.research.google.com</a></p><pre><code>  
# 2. vLLM - Fast-API based server for Online Serving
# OpenAI API-compatible server

# Server
! python -m vllm.entrypoints.openai.api_server 
--host 127.0.0.1 
--port 8888 
--model meta-llama/Llama-2-7b

# Client 
!curl http://127.0.0.1:8888/v1/completions 
-H "Content-Type: application/json" 
-d '{
  "model": "meta-llama/Llama-2-7b",
  "prompt": "Paris is a",
  "max_tokens": 7,
  "temperature": 0
  }'</code></pre><p>LMSYS introduced <a href="https://lmsys.org/blog/2023-03-30-vicuna/">Vicuna</a> chatbot models, which are now used by millions in <a href="https://arena.lmsys.org/">Chatbot Arena</a>. Initially, <a href="https://github.com/lm-sys/FastChat/blob/main/fastchat/serve/vllm_worker.py">FastChat</a> used HF Transformers for serving, but as traffic surged, vLLM was integrated to handle up to 5x more traffic, significantly improving throughput by up to 30x over the initial HF backend.</p><h3>Conclusion</h3><p>vLLM, powered by PagedAttention, represents a significant advancement in LLM serving, offering a solution that is not only fast and efficient but also cost-effective. Its practical applications, such as serving Vicuna chatbot models, showcase its potential to revolutionize how LLMs are used across various industries. With vLLM, the future of LLM serving looks promising, offering high throughput and low latency without compromising on performance</p><blockquote><p><em><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p></blockquote><blockquote><p><em><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></em></p></blockquote><blockquote><p><em><strong>Visit my technical channel on</strong> <strong><a href="http://www.youtube.com/@aboniasojasingarayar3097">Youtube</a></strong></em></p></blockquote><blockquote><p>Thanks for&nbsp;Reading!</p></blockquote>]]></content:encoded></item><item><title><![CDATA[Summarization with LangChain]]></title><description><![CDATA[Stuff &#8212; Map_reduce &#8212; Refine]]></description><link>https://aboniasojasingarayar.substack.com/p/summarization-with-langchain-b3d83c030889</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/summarization-with-langchain-b3d83c030889</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Tue, 23 Apr 2024 08:37:52 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/725b3392-64fc-4465-b421-bc7b6b0cdda9_800x450.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4><code>Stuff&#8202;&#8212;&#8202;Map_reduce&#8202;&#8212;&#8202;Refine</code></h4><p>I&nbsp;<br>&nbsp;recently wrapped a tutorial on summarization techniques in LangChain. This article covers the basic usage of document summarization techniques and provides insights into various summarization methods. Additionally, to learn more and to explore how to validate intermediate results from the output of each of these techniques. For those interested in delving deeper into this topic, I invite you to explore the complete and comprehensive tutorial available <strong><a href="https://youtu.be/w6wOhSThnoo">here</a></strong>.</p><p>Explore daily 5-minute articles featuring the newest breakthroughs, research, models, and repositories.Subscribe to my <strong><a href="https://abonia1.github.io/newsletter/">Newsletter</a></strong></p><p><strong><a href="https://abonia1.github.io/newsletter/" title="https://abonia1.github.io/newsletter/">Newsletter</a></strong><a href="https://abonia1.github.io/newsletter/" title="https://abonia1.github.io/newsletter/"><br></a><em><a href="https://abonia1.github.io/newsletter/" title="https://abonia1.github.io/newsletter/">Abonia&#8202;&#8212;&#8202;ML Engineer</a></em><a href="https://abonia1.github.io/newsletter/" title="https://abonia1.github.io/newsletter/">abonia1.github.io</a></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6A_N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6A_N!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!6A_N!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!6A_N!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!6A_N!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6A_N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6A_N!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!6A_N!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!6A_N!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!6A_N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F795bcbbb-ae18-4979-96af-4f3ba75258f9_800x450.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Link to Tutorial&#8202;&#8212;&#8202;<strong><a href="https://youtu.be/w6wOhSThnoo">https://youtu.be/w6wOhSThnoo</a></strong></figcaption></figure></div><p>Summarization is a critical aspect of natural language processing (NLP), enabling the condensation of large volumes of text into concise summaries. LangChain, a powerful tool in the NLP domain, offers three distinct summarization techniques: <code>stuff</code>, <code>map_reduce</code>, and <code>refine</code>. Each method has its unique advantages and limitations, making them suitable for different scenarios. This article delves into the details of these techniques, their pros and cons, and the ideal scenarios for their application.</p><p>Complete implementation code and data used in the tutorial available in the below <a href="https://github.com/Abonia1/Langchain-Summarizer">repository</a>.</p><p><strong><a href="https://github.com/Abonia1/Langchain-Summarizer" title="https://github.com/Abonia1/Langchain-Summarizer">GitHub - Abonia1/Langchain-Summarizer</a></strong><a href="https://github.com/Abonia1/Langchain-Summarizer" title="https://github.com/Abonia1/Langchain-Summarizer"><br></a><em><a href="https://github.com/Abonia1/Langchain-Summarizer" title="https://github.com/Abonia1/Langchain-Summarizer">Contribute to Abonia1/Langchain-Summarizer development by creating an account on GitHub.</a></em><a href="https://github.com/Abonia1/Langchain-Summarizer" title="https://github.com/Abonia1/Langchain-Summarizer">github.com</a></p><h3>Summarization Techniques</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c8Wk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c8Wk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png 424w, https://substackcdn.com/image/fetch/$s_!c8Wk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png 848w, https://substackcdn.com/image/fetch/$s_!c8Wk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png 1272w, https://substackcdn.com/image/fetch/$s_!c8Wk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c8Wk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c8Wk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png 424w, https://substackcdn.com/image/fetch/$s_!c8Wk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png 848w, https://substackcdn.com/image/fetch/$s_!c8Wk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png 1272w, https://substackcdn.com/image/fetch/$s_!c8Wk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fde106-4e3f-49ff-b5cb-c8458333676f_800x356.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Courtesy of langchain</figcaption></figure></div><p><strong><a href="https://python.langchain.com/docs/use_cases/summarization" title="https://python.langchain.com/docs/use_cases/summarization">Summarization | &#129436;&#65039;&#128279; Langchain</a></strong><a href="https://python.langchain.com/docs/use_cases/summarization" title="https://python.langchain.com/docs/use_cases/summarization"><br></a><em><a href="https://python.langchain.com/docs/use_cases/summarization" title="https://python.langchain.com/docs/use_cases/summarization">Open In Colab</a></em><a href="https://python.langchain.com/docs/use_cases/summarization" title="https://python.langchain.com/docs/use_cases/summarization">python.langchain.com</a></p><h3>1. Stuff&nbsp;Chain</h3><p>The <code>stuff</code> chain is particularly effective for handling large documents. It works by converting the document into smaller chunks, processing each chunk individually, and then combining the summaries to generate a final summary. This method is ideal for managing huge files and can be facilitated with the help of a recursive character text splitter.</p><p><strong>Pros:</strong></p><ul><li><p>Efficiently handles large documents.</p></li><li><p>Allows for chunk-wise summarization, making it suitable for managing huge files.</p></li></ul><p><strong>Cons:</strong></p><ul><li><p>LLM context window</p></li></ul><pre><code>from langchain.chains.combine_documents.stuff import StuffDocumentsChain
from langchain.chains.llm import LLMChain
from langchain.prompts import PromptTemplate

# Define prompt
prompt_template = """Write a concise summary of the following:
"{text}"
CONCISE SUMMARY:"""
prompt = PromptTemplate.from_template(prompt_template)
# Define LLM chain
llm = ChatOpenAI(temperature=0, model_name="gpt-3.5-turbo-16k")
llm_chain = LLMChain(llm=llm, prompt=prompt)
# Define StuffDocumentsChain
stuff_chain = StuffDocumentsChain(llm_chain=llm_chain, document_variable_name="text")
docs = loader.load()
print(stuff_chain.run(docs))</code></pre><h3>Map-Reduce Method</h3><p>The Map-Reduce method involves summarizing each document individually (map step) and then combining these summaries into a final summary (reduce step). This approach is more scalable and can handle larger volumes of text.The <code>map_reduce</code> technique is designed for summarizing large documents that exceed the token limit of the language model. It involves dividing the document into chunks, generating summaries for each chunk, and then combining these summaries to create a final summary. This method is efficient for handling large files and significantly reduces processing time.</p><p><strong>Pros:</strong></p><ul><li><p>Effectively handles large documents by dividing them into manageable chunks.</p></li><li><p>Reduces processing time by processing chunks individually.</p></li></ul><p><strong>Cons:</strong></p><ul><li><p>Requires extra steps in combining individual summaries, which can add complexity to the process.</p></li></ul><p>Here&#8217;s an example of how to implement the Map-Reduce method:</p><pre><code>from langchain.chains import MapReduceDocumentsChain, ReduceDocumentsChain
from langchain_text_splitters import CharacterTextSplitter

# Map
map_template = """The following is a set of documents
{docs}
Based on this list of docs, please identify the main themes 
Helpful Answer:"""
map_prompt = PromptTemplate.from_template(map_template)
map_chain = LLMChain(llm=llm, prompt=map_prompt)
# Reduce
reduce_template = """The following is set of summaries:
{docs}
Take these and distill it into a final, consolidated summary of the main themes. 
Helpful Answer:"""
reduce_prompt = PromptTemplate.from_template(reduce_template)
reduce_chain = LLMChain(llm=llm, prompt=reduce_prompt)
# Combine documents by mapping a chain over them, then combining results
map_reduce_chain = MapReduceDocumentsChain(
    llm_chain=map_chain,
    reduce_documents_chain=reduce_documents_chain,
    document_variable_name="docs",
    return_intermediate_steps=False,
)
text_splitter = CharacterTextSplitter.from_tiktoken_encoder(chunk_size=1000, chunk_overlap=0)
split_docs = text_splitter.split_documents(docs)
print(map_reduce_chain.run(split_docs))</code></pre><h3>Refine Method</h3><p>The Refine method iteratively updates its answer by looping over the input documents. For each document, it passes all non-document inputs, the current document, and the latest intermediate answer to an LLM chain to get a new answer. This method is useful for refining a summary based on new context.The <code>refine</code> technique is a simpler alternative to the <code>map_reduce</code> technique. It involves generating a summary for the first chunk, combining it with the second chunk, generating another summary, and continuing this process until a final summary is achieved. This method is suitable for large documents but requires less complexity compared to <code>map_reduce</code>.</p><p><strong>Pros:</strong></p><ul><li><p>Simpler than the <code>map_reduce</code> technique.</p></li><li><p>Achieves similar results for large documents with less complexity.</p></li></ul><p><strong>Cons:</strong></p><ul><li><p>Limited functionality compared to other techniques.</p></li></ul><p>Here&#8217;s how you can implement the Refine method:</p><pre><code>from langchain.chains.summarize import load_summarize_chain
prompt = """
                  Please provide a summary of the following text.
                  TEXT: {text}
                  SUMMARY:
                  """

question_prompt = PromptTemplate(
    template=question_prompt_template, input_variables=["text"]
)

refine_prompt_template = """
              Write a concise summary of the following text delimited by triple backquotes.
              Return your response in bullet points which covers the key points of the text.
              ```{text}```
              BULLET POINT SUMMARY:
              """

refine_template = PromptTemplate(
    template=refine_prompt_template, input_variables=["text"]

# Load refine chain
chain = load_summarize_chain(
    llm=llm,
    chain_type="refine",
    question_prompt=question_prompt,
    refine_prompt=refine_prompt,
    return_intermediate_steps=True,
    input_key="input_documents",
    output_key="output_text",
)
result = chain({"input_documents": split_docs}, return_only_outputs=True)</code></pre><h3>Choosing the Right Technique</h3><p>The choice of summarization technique depends on the specific requirements of the task at hand. For large documents, the <code>map_reduce</code> and <code>refine</code> techniques are recommended due to their ability to handle chunk-wise summarization efficiently. The <code>stuff</code> chain is particularly useful for documents that are too large to be processed in a single go, offering a practical solution for managing huge files.</p><p>Each method has its advantages and is suitable for different scenarios. The Stuff method is straightforward but may not scale well with large volumes of text. The Map-Reduce method is more scalable and can handle larger documents but requires more setup. The Refine method is useful for iteratively refining a summary based on new context, making it a good choice for dynamic summarization tasks.</p><blockquote><p><em><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p></blockquote><blockquote><p><em><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></em></p></blockquote><blockquote><p><em><strong>Visit my technical channel on</strong> <strong><a href="http://www.youtube.com/@aboniasojasingarayar3097">Youtube</a></strong></em></p></blockquote><blockquote><p><em>Thanks for&nbsp;Reading!</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Books on LLM and NLP in 2024]]></title><description><![CDATA[Exploring the Best Books for Understanding and Implementing LLMs in NLP]]></description><link>https://aboniasojasingarayar.substack.com/p/books-on-llm-and-nlp-in-2024-aea05617d557</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/books-on-llm-and-nlp-in-2024-aea05617d557</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Wed, 17 Apr 2024 07:32:05 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8c8836e8-64a1-4b6c-8b8d-a1ccf732fcbc_800x450.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Exploring the Best Books for Understanding and Implementing LLMs in&nbsp;NLP</h4><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eGzk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eGzk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!eGzk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!eGzk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!eGzk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eGzk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eGzk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!eGzk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!eGzk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!eGzk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff3d02483-e7a7-4166-9e87-ebfa3c266092_800x450.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Image by&nbsp;Author</figcaption></figure></div><p>Large Language Models (LLMs) have significantly transformed the landscape of Natural Language Processing (NLP), offering precise and efficient methods for understanding and generating human language. These models are now integral to various applications, including chatbots, language translation, text summarization, and sentiment analysis, across numerous industries. However, mastering LLMs can be challenging due to their complexity and the sophisticated algorithms behind them. To aid in this learning journey, a plethora of resources, including books, crash courses, university courses, blogs, and papers, have been developed. This article will delve into some of the top books that provide valuable insights and practical knowledge for working with LLMs, catering to both beginners and experienced practitioners in the field.</p><p>Explore in below <a href="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop">link</a> to find out more about other favorite NLP and LLM books and and discover why they&#8217;re highly regarded.</p><p><strong><a href="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop" title="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop">Abonia Sojasingarayar on LinkedIn: &#128218; Books on LLM &amp;amp; NLP 2024 &#128218; &#128312; Practical Natural Language&#8230;</a></strong><a href="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop" title="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop"><br></a><em><a href="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop" title="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop">&#128218; Books on LLM &amp;amp; NLP 2024 &#128218; &#128312; Practical Natural Language Processing - O&amp;#39;Reilly: https://lnkd.in/eKBHvdzM &#128312;&#8230;</a></em><a href="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop" title="https://www.linkedin.com/posts/aboniasojasingarayar_books-on-llm-nlp-2024-practical-activity-7176132889896452096-UhBM?utm_source=share&amp;utm_medium=member_desktop">www.linkedin.com</a></p><blockquote><p>Let me know in the comment here&nbsp;, whats your favorite NLP and LLM books(within or beyond the below list) and why you like it&nbsp;please?</p></blockquote><h3>Curated list</h3><p>&#128312; Practical Natural Language Processing&#8202;&#8212;&#8202;O&#8217;Reilly: <a href="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/">https://www.oreilly.com/library/view/practical-natural-language/9781492054047/</a></p><p>&#128312;Natural Language Processing with Transformers&#8202;&#8212;&#8202;O&#8217;Reilly&nbsp;: <a href="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/">https://www.oreilly.com/library/view/natural-language-processing/9781098136789/</a></p><p>&#128312;Transformers for Natural Language Processing&#8202;&#8212;&#8202;Packt: <a href="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791">https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791</a></p><p>&#128312;GPT-3: Building Innovative NLP Products Using Large Language Models&#8202;&#8212;&#8202;O&#8217;Reilly: <a href="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624">https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624</a></p><p>&#128312;Hands-On Generative AI with Transformers and Diffusion Models&#8202;&#8212;&#8202;O&#8217;Reilly:https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/</p><p>&#128312;Quick Start Guide to Large Language Models-Strategies and Best Practices for Using ChatGPT and Other LLMs&#8202;&#8212;&#8202;O&#8217;Reilly: <a href="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/">https://www.oreilly.com/library/view/quick-start-guide/9780138199425/</a></p><p>&#128312;Understanding Large Language Models: Learning Their Underlying Concepts and Technologies: <a href="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE">https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE</a></p><p>&#128312;Generative AI with LangChain: Build large language model (LLM) apps with Python, ChatGPT, and other LLMs&nbsp;: <a href="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463">https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463</a></p><p>&#128312;Generative AI on AWS&#8202;&#8212;&#8202;O&#8217;Reilly: <a href="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/">https://www.oreilly.com/library/view/generative-ai-on/9781098159214/</a></p><p>&#128312;Build a Large Language Model (From Scratch): <a href="https://www.manning.com/books/build-a-large-language-model-from-scratch">https://www.manning.com/books/build-a-large-language-model-from-scratch</a></p><p>&#128312;Designing Large Language Model Applications&#8202;&#8212;&#8202;O&#8217;Reilly: <a href="https://www.oreilly.com/library/view/designing-large-language/9781098150495/">https://www.oreilly.com/library/view/designing-large-language/9781098150495/</a></p><p>&#128312;Pretrain Vision and Large Language Models in Python&nbsp;: <a href="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/">https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/</a></p><h4>1. Practical Natural Language Processing&#8202;&#8212;&#8202;O&#8217;Reilly</h4><p>Authored by Sowmya Vajjala, Bodhisattwa Majumder, Anuj Gupta, and Harshit Surana, this book is a comprehensive guide to building, iterating, and scaling NLP systems in a business setting. It covers a wide spectrum of problem statements, tasks, and solution approaches within NLP, emphasizing the importance of adapting solutions for different industry verticals such as healthcare, social media, and retail. The book is designed to help software engineers and data scientists navigate the complexities of NLP, offering practical advice on best practices for release, deployment, and DevOps for NLP systems.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_4yo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_4yo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg 424w, https://substackcdn.com/image/fetch/$s_!_4yo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg 848w, https://substackcdn.com/image/fetch/$s_!_4yo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!_4yo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_4yo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_4yo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg 424w, https://substackcdn.com/image/fetch/$s_!_4yo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg 848w, https://substackcdn.com/image/fetch/$s_!_4yo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!_4yo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49550665-7849-4e80-8a25-4ffbabc07b7a_762x1000.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/" title="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/">Practical Natural Language Processing</a></strong><a href="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/" title="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/"><br></a><em><a href="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/" title="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/">Many books and courses tackle natural language processing (NLP) problems with toy use cases and well-defined datasets&#8230;</a></em><a href="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/" title="https://www.oreilly.com/library/view/practical-natural-language/9781492054047/">www.oreilly.com</a></p><h4>2. Natural Language Processing with Transformers&#8202;&#8212;&#8202;O&#8217;Reilly</h4><p>This book, written by Lewis Tunstall, Leandro von Werra, and Thomas Wolf, provides an in-depth look at transformers, the dominant architecture for achieving state-of-the-art results in NLP. Since their introduction in 2017, transformers have revolutionized the field, offering significant improvements over previous models. The book is a valuable resource for understanding the underlying mechanisms of transformers and how they can be applied to various NLP tasks. It covers the practical aspects of training and scaling these large models using Hugging Face Transformers, a Python-based deep learning library, and offers insights into real-world applications of transformers, such as writing realistic news stories and creating chatbots.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IJ3e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IJ3e!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg 424w, https://substackcdn.com/image/fetch/$s_!IJ3e!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg 848w, https://substackcdn.com/image/fetch/$s_!IJ3e!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!IJ3e!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IJ3e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!IJ3e!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg 424w, https://substackcdn.com/image/fetch/$s_!IJ3e!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg 848w, https://substackcdn.com/image/fetch/$s_!IJ3e!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!IJ3e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6df65022-b13b-4772-813c-b263b53ed090_381x500.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/" title="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/">Natural Language Processing with Transformers, Revised Edition</a></strong><a href="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/" title="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/"><br></a><em><a href="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/" title="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/">Since their introduction in 2017, transformers have quickly become the dominant architecture for achieving&#8230;</a></em><a href="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/" title="https://www.oreilly.com/library/view/natural-language-processing/9781098136789/">www.oreilly.com</a></p><h4>3. Transformers for Natural Language Processing&#8202;&#8212;&#8202;Packt</h4><p>This book offers a deep dive into the world of transformers, focusing on their application in NLP. It covers the core concepts of transformers, including self-attention mechanisms, and explores how these models can be used to enhance language modeling capabilities. The book is an excellent resource for those looking to understand the technical aspects of transformers and their role in NLP. It provides a comprehensive guide to transformer architectures, starting with the original Transformer, before moving on to RoBERTa, BERT, and DistilBERT models. The book also covers training methods for smaller Transformers that can outperform GPT-3 in some cases and advanced language understanding techniques such as optimizing social network datasets and fake news identification.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!O00d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!O00d!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png 424w, https://substackcdn.com/image/fetch/$s_!O00d!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png 848w, https://substackcdn.com/image/fetch/$s_!O00d!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png 1272w, https://substackcdn.com/image/fetch/$s_!O00d!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!O00d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!O00d!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png 424w, https://substackcdn.com/image/fetch/$s_!O00d!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png 848w, https://substackcdn.com/image/fetch/$s_!O00d!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png 1272w, https://substackcdn.com/image/fetch/$s_!O00d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4435a69b-e498-4b86-b4f5-84cd5e596700_800x988.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791" title="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791">Transformers for Natural Language Processing | Packt</a></strong><a href="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791" title="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791"><br></a><em><a href="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791" title="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791">Publisher's Note: A new edition of this book is out now that includes working with GPT-3 and comparing the results with&#8230;</a></em><a href="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791" title="https://www.packtpub.com/product/transformers-for-natural-language-processing/9781800565791">www.packtpub.com</a></p><h4>4. GPT-3: Building Innovative NLP Products Using Large Language Models&#8202;&#8212;&#8202;O&#8217;Reilly</h4><p>This book delves into the capabilities of GPT-3, one of the most advanced LLMs available today. It provides insights into building innovative NLP products using GPT-3, covering topics such as model fine-tuning, retrieval-augmented generation, and reinforcement learning from human feedback. The book is a practical guide for developers and businesses looking to leverage the power of GPT-3 for their applications.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MdY0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MdY0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg 424w, https://substackcdn.com/image/fetch/$s_!MdY0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg 848w, https://substackcdn.com/image/fetch/$s_!MdY0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!MdY0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MdY0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MdY0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg 424w, https://substackcdn.com/image/fetch/$s_!MdY0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg 848w, https://substackcdn.com/image/fetch/$s_!MdY0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!MdY0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81c78e27-55f0-4a2f-88e0-6c6ddcd2e3bb_731x1000.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624" title="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624">GPT-3: Building Innovative NLP Products Using Large Language Models</a></strong><a href="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624" title="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624"><br></a><em><a href="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624" title="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624">GPT-3: NLP with LLMs is a unique, pragmatic take on Generative Pre-trained Transformer 3, the famous AI language model&#8230;</a></em><a href="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624" title="https://www.amazon.com/GPT-3-Building-Innovative-Products-Language/dp/1098113624">www.amazon.com</a></p><h4>5. Hands-On Generative AI with Transformers and Diffusion Models&#8202;&#8212;&#8202;O&#8217;Reilly</h4><p>This book is a hands-on guide to generative AI, focusing on transformers and diffusion models. It covers the generative AI project lifecycle, including use case definition, model selection, fine-tuning, and deployment. The book is designed to help readers apply generative AI to their business use cases, offering practical advice on model selection, fine-tuning, and integration with existing software ecosystems.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bVX3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bVX3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bVX3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bVX3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bVX3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bVX3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bVX3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg 424w, https://substackcdn.com/image/fetch/$s_!bVX3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg 848w, https://substackcdn.com/image/fetch/$s_!bVX3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!bVX3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6bc41f49-fc9b-4b47-bf38-d64631d1c791_800x1049.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/" title="https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/">Hands-On Generative AI with Transformers and Diffusion Models</a></strong><a href="https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/" title="https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/"><br></a><em><a href="https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/" title="https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/">Learn how to use generative media techniques with AI to create novel images or music in this practical, hands-on guide&#8230;</a></em><a href="https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/" title="https://www.oreilly.com/library/view/hands-on-generative-ai/9781098149239/">www.oreilly.com</a></p><h4>6. Quick Start Guide to Large Language Models-Strategies and Best Practices for Using ChatGPT and Other LLMs&#8202;&#8212;&#8202;O&#8217;Reilly</h4><p>This guide provides a quick start to using large language models, with a focus on ChatGPT and other LLMs. It offers strategies and best practices for implementing LLMs in projects, covering topics such as model selection, fine-tuning, and deployment. The book is an excellent resource for developers and businesses looking to quickly integrate LLMs into their applications.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ffsj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ffsj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ffsj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ffsj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ffsj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ffsj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ffsj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ffsj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ffsj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ffsj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2645bc3-ce5f-4807-92e3-bd3d15ba0b4e_765x1000.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/" title="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/">Quick Start Guide to Large Language Models: Strategies and Best Practices for Using ChatGPT and&#8230;</a></strong><a href="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/" title="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/"><br></a><em><a href="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/" title="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/">The advancement of Large Language Models (LLMs) has revolutionized the field of Natural Language Processing in recent&#8230;</a></em><a href="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/" title="https://www.oreilly.com/library/view/quick-start-guide/9780138199425/">www.oreilly.com</a></p><h4>7. Understanding Large Language Models: Learning Their Underlying Concepts and Technologies</h4><p>Authored by Thimira Amaratunga, this book offers a comprehensive understanding of LLMs, covering their underlying concepts and technologies. It explores the rise of conversational AIs, the evolution of NLP, and the unique capabilities of LLMs. The book is designed to equip readers with the knowledge to implement LLMs in their projects, offering insights into the architectures of popular LLMs and the opportunities they present.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-NzG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-NzG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-NzG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-NzG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-NzG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-NzG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-NzG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-NzG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-NzG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-NzG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F570f5daf-f84c-43c8-a489-c5033905ba3e_659x1000.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE" title="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE">Understanding Large Language Models</a></strong><a href="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE" title="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE"><br></a><em><a href="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE" title="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE">This book will teach you the underlying concepts of large language models (LLMs), as well as the technologies&#8230;</a></em><a href="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE" title="https://link.springer.com/book/10.1007/979-8-8688-0017-7?source=shoppingads&amp;locale=en-fr&amp;gad_source=1&amp;gclid=CjwKCAiAopuvBhBCEiwAm8jaMZV12rj1v0DL_ZMWo1qwtLlsrbmXbUrU3pqPbEk31Kb7-P3NsPCV4RoCtesQAvD_BwE">link.springer.com</a></p><h4>8. Generative AI with LangChain: Build large language model (LLM) apps with Python, ChatGPT, and other&nbsp;LLMs</h4><p>This book provides a practical guide to building LLM applications using LangChain, Python, and ChatGPT. It covers the generative AI project lifecycle, including model selection, fine-tuning, and deployment. The book is an excellent resource for developers looking to leverage generative AI for their applications, offering practical advice on model selection, fine-tuning, and integration with existing software ecosystems.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rmql!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rmql!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg 424w, https://substackcdn.com/image/fetch/$s_!rmql!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg 848w, https://substackcdn.com/image/fetch/$s_!rmql!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!rmql!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rmql!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rmql!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg 424w, https://substackcdn.com/image/fetch/$s_!rmql!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg 848w, https://substackcdn.com/image/fetch/$s_!rmql!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!rmql!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F165c9955-8c06-48fd-9f4b-4687a867080f_800x986.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463" title="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463">Generative AI with LangChain: Build large language model (LLM) apps with Python, ChatGPT and other&#8230;</a></strong><a href="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463" title="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463"><br></a><em><a href="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463" title="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463">Not&#233; /5. Retrouvez Generative AI with LangChain: Build large language model (LLM) apps with Python, ChatGPT and other&#8230;</a></em><a href="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463" title="https://www.amazon.fr/Generative-AI-LangChain-language-ChatGPT/dp/1835083463">www.amazon.fr</a></p><h4>9. Generative AI on AWS&#8202;&#8212;&#8202;O&#8217;Reilly</h4><p>This book, authored by Chris Fregly, Antje Barth, and Shelbee Eigenbrode, offers a comprehensive guide to applying generative AI on AWS. It covers the generative AI project lifecycle, including model selection, fine-tuning, and deployment. The book is designed to help readers apply generative AI to their business use cases, offering practical advice on model selection, fine-tuning, and integration with existing software ecosystems.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LFqD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LFqD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LFqD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LFqD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LFqD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LFqD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LFqD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LFqD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LFqD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LFqD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F293cf7a1-477b-4e86-bfea-b550eae144f2_381x500.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/" title="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/">Generative AI on AWS</a></strong><a href="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/" title="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/"><br></a><em><a href="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/" title="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/">Companies today are moving rapidly to integrate generative AI into their products and services. But there's a great&#8230;</a></em><a href="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/" title="https://www.oreilly.com/library/view/generative-ai-on/9781098159214/">www.oreilly.com</a></p><h4>10. Build a Large Language Model (From&nbsp;Scratch)</h4><p>This book provides a step-by-step guide to building a large language model from scratch. It covers the technical aspects of building LLMs, including model architecture, training, and deployment. The book is an excellent resource for developers and researchers looking to build their own LLMs, offering practical advice on model architecture, training, and deployment.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!W7rL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!W7rL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg 424w, https://substackcdn.com/image/fetch/$s_!W7rL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg 848w, https://substackcdn.com/image/fetch/$s_!W7rL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!W7rL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!W7rL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!W7rL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg 424w, https://substackcdn.com/image/fetch/$s_!W7rL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg 848w, https://substackcdn.com/image/fetch/$s_!W7rL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!W7rL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad4e9a5e-a4d9-4b08-a25b-0eb42bc3edc4_600x752.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.manning.com/books/build-a-large-language-model-from-scratch" title="https://www.manning.com/books/build-a-large-language-model-from-scratch">Build a Large Language Model (From Scratch)</a></strong><a href="https://www.manning.com/books/build-a-large-language-model-from-scratch" title="https://www.manning.com/books/build-a-large-language-model-from-scratch"><br></a><em><a href="https://www.manning.com/books/build-a-large-language-model-from-scratch" title="https://www.manning.com/books/build-a-large-language-model-from-scratch">Learn how to create, train, and tweak large language models (LLMs) by building one from the ground up! In Build a Large&#8230;</a></em><a href="https://www.manning.com/books/build-a-large-language-model-from-scratch" title="https://www.manning.com/books/build-a-large-language-model-from-scratch">www.manning.com</a></p><h4>11. Designing Large Language Model Applications&#8202;&#8212;&#8202;O&#8217;Reilly</h4><p>Authored by Suhas Pai, this book offers practical advice on building useful products that incorporate the power of language models. It covers the tools, techniques, and playbooks for transitioning from demos and prototypes to full-fledged applications. The book is designed to help readers develop an intuition about the Transformer architecture and the impact of each architectural decision.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tc9M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tc9M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tc9M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tc9M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tc9M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tc9M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tc9M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg 424w, https://substackcdn.com/image/fetch/$s_!tc9M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg 848w, https://substackcdn.com/image/fetch/$s_!tc9M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!tc9M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe1b8d876-caa6-4761-aa8b-b77c88399b91_400x525.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.oreilly.com/library/view/designing-large-language/9781098150495/" title="https://www.oreilly.com/library/view/designing-large-language/9781098150495/">Designing Large Language Model Applications</a></strong><a href="https://www.oreilly.com/library/view/designing-large-language/9781098150495/" title="https://www.oreilly.com/library/view/designing-large-language/9781098150495/"><br></a><em><a href="https://www.oreilly.com/library/view/designing-large-language/9781098150495/" title="https://www.oreilly.com/library/view/designing-large-language/9781098150495/">Transformer-based language models are powerful tools for solving a variety of language tasks and represent a phase&#8230;</a></em><a href="https://www.oreilly.com/library/view/designing-large-language/9781098150495/" title="https://www.oreilly.com/library/view/designing-large-language/9781098150495/">www.oreilly.com</a></p><h4>12. Pretrain Vision and Large Language Models in&nbsp;Python</h4><p>This book provides a comprehensive guide to pretraining vision and large language models in Python. It covers the technical aspects of pretraining models, including model architecture, training, and deployment. The book is an excellent resource for developers and researchers looking to pretrain their own models, offering practical advice on model architecture, training, and deployment.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RNit!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RNit!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg 424w, https://substackcdn.com/image/fetch/$s_!RNit!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg 848w, https://substackcdn.com/image/fetch/$s_!RNit!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!RNit!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RNit!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RNit!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg 424w, https://substackcdn.com/image/fetch/$s_!RNit!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg 848w, https://substackcdn.com/image/fetch/$s_!RNit!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!RNit!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe78e9af9-4934-49b1-be1b-d32ba5e8ce3c_405x500.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong><a href="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/" title="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/">Pretrain Vision and Large Language Models in Python</a></strong><a href="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/" title="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/"><br></a><em><a href="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/" title="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/">Master the art of training vision and large language models with conceptual fundaments and industry-expert guidance&#8230;</a></em><a href="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/" title="https://www.oreilly.com/library/view/pretrain-vision-and/9781804618257/">www.oreilly.com</a></p><h4>Conclusion</h4><p>In conclusion, these books offer a wide range of perspectives on NLP and LLMs, from practical applications to theoretical foundations. Whether you&#8217;re a beginner looking to get started in the field or an experienced practitioner seeking to deepen your knowledge, these resources provide valuable insights and practical advice to help you navigate the complexities of NLP and LLMs.</p><p><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></p><p><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></p><p><strong>Visit my technical channel on</strong> <strong><a href="https://www.youtube.com/channel/UCGphGM_oeR4r9dqVs71Jc5w">Youtube</a></strong></p><blockquote><p>Thanks for&nbsp;Reading!</p></blockquote>]]></content:encoded></item><item><title><![CDATA[Deploying a RAG Application in AWS Lambda using Docker and ECR]]></title><description><![CDATA[Lambda &#8212; ECR &#8212; Docker &#8212; LangChain &#8212; OpenAI]]></description><link>https://aboniasojasingarayar.substack.com/p/deploying-a-rag-application-in-aws-lambda-using-docker-and-ecr-08e246a7c515</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/deploying-a-rag-application-in-aws-lambda-using-docker-and-ecr-08e246a7c515</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Tue, 02 Apr 2024 07:48:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/gicsb9p7uj4" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Lambda&#8202;&#8212;&#8202;ECR&#8202;&#8212;&#8202;Docker&#8202;&#8212;&#8202;LangChain&#8202;&#8212;&#8202;OpenAI</h4><p>Deploying a RAG (Retrieval-Augmented Generation) application in AWS Lambda using Docker and Amazon Elastic Container Registry (ECR) with LangChain involves several steps and services. This article will explain each service and how they work together in an integrated way, followed by the steps to deploy it.</p><div class="captioned-image-container"><figure><div id="youtube2-gicsb9p7uj4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;gicsb9p7uj4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/gicsb9p7uj4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><figcaption class="image-caption">Build and Deploy RAG in AWS</figcaption></figure></div><p>Here is the link to the complete tutorial on <a href="https://youtu.be/gicsb9p7uj4?si=9F2l6z1rNpkUOoIR">D</a><strong><a href="https://youtu.be/gicsb9p7uj4?si=9F2l6z1rNpkUOoIR">eploying RAG in AWS</a></strong>.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ua_C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ua_C!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!ua_C!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!ua_C!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!ua_C!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ua_C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ua_C!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!ua_C!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!ua_C!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!ua_C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbc728ba-e20c-488d-bb0b-a2e41e89ac43_800x450.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Image by&nbsp;Author</figcaption></figure></div><p>Before proceeding further, I would like to kindly suggest that you take some time to read and watch the tutorial titled &#8220;<strong><a href="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749">Build and Deploy LLM Application in AWS</a></strong>.&#8221; This tutorial can provide you with a solid foundation on Lambda LLM application deployment.</p><p><strong><a href="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749" title="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749">Build and Deploy LLM Application in AWS</a></strong><a href="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749" title="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749"><br></a><em><a href="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749" title="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749">AWS Lambda&#8202;&#8212;&#8202;BedRock&#8202;&#8212;&#8202;LangChain</a></em><a href="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749" title="https://medium.com/@abonia/build-and-deploy-llm-application-in-aws-cca46c662749">medium.com</a></p><h3>Services Overview</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BzU2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BzU2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!BzU2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!BzU2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!BzU2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BzU2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BzU2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!BzU2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!BzU2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!BzU2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717a9eb3-7d3e-4481-bf99-b0214e0b6904_800x450.png 1456w" sizes="100vw"></picture><div></div></div></a><figcaption class="image-caption">Architecture Overview&#8202;&#8212;&#8202;Image by&nbsp;Author</figcaption></figure></div><p><a href="https://aws.amazon.com/lambda/getting-started/?gclid=CjwKCAjwzN-vBhAkEiwAYiO7oF7Q4aEK3hKUA0HVPEe3wXKen_tVUDmxun4s5PjasxZ-deFf-j19vxoCovkQAvD_BwE&amp;trk=65546593-a75a-4317-b5dc-80218abfdb10&amp;sc_channel=ps&amp;s_kwcid=AL!4422!3!651542249938!e!!g!!aws%20lambda&amp;ef_id=CjwKCAjwzN-vBhAkEiwAYiO7oF7Q4aEK3hKUA0HVPEe3wXKen_tVUDmxun4s5PjasxZ-deFf-j19vxoCovkQAvD_BwE:G:s&amp;s_kwcid=AL!4422!3!651542249938!e!!g!!aws%20lambda!19835810591!150095231954">AWS Lambda</a>: A serverless compute service that runs your code in response to events and automatically manages the underlying compute resources for you. It allows you to run code without provisioning or managing servers.</p><p>-<a href="https://aws.amazon.com/ecr/"> Amazon ECR</a>: A fully-managed container registry that makes it easy for developers to store, manage, and deploy Docker container images. It&#8217;s integrated with Amazon ECS and AWS Fargate, simplifying your development to production workflow.</p><p>- <a href="https://www.docker.com/">Docker</a>: A platform that uses containerization to package up an application with all of the parts it needs, such as libraries and other dependencies, and ship it all out as one package. This ensures that the application runs quickly and reliably on any server.</p><blockquote><p>Dowload Link: <a href="https://www.docker.com/products/docker-desktop/">https://www.docker.com/products/docker-desktop/</a></p></blockquote><p>- <a href="https://www.langchain.com/">LangChain</a>: A framework for building and deploying language models, providing tools for document loading, vector storage, embeddings, and more. It&#8217;s used here to facilitate the deployment of the RAG model.</p><h3>Integration and Deployment Steps</h3><ol><li><p><strong>Prepare Your Environment: </strong>Ensure you have the AWS CLI installed and configured, Docker installed, and Python 3.11 or later installed and configured. You&#8217;ll also need an active AWS account with the necessary permissions.</p></li></ol><blockquote><p>Download Link: <a href="https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html">https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html</a></p></blockquote><p>2. <strong>Create a Dockerfile:</strong> This file defines the environment in which your Lambda function will run. It starts from a base image provided by AWS (<em>`public.ecr.aws/lambda/python:3.12`</em>) and copies your application code and dependencies into the image. It also installs any necessary Python packages from a `requirements.txt` file as below:</p><pre><code>langchain_community
boto3==1.34.37
numpy
langchain
langchainhub
langchain-openai
chromadb
bs4
tiktoken
openai</code></pre><pre><code># Use the AWS base image for Python 3.12
FROM public.ecr.aws/lambda/python:3.12

# Install build-essential to get the C++ compiler and other necessary tools
RUN microdnf update -y &amp;&amp; microdnf install -y gcc-c++ make

# Copy requirements.txt
COPY requirements.txt ${LAMBDA_TASK_ROOT}

# Install the specified packages
RUN pip install -r requirements.txt

# Copy function code
COPY lambda_function.py ${LAMBDA_TASK_ROOT}

# Set the permissions to make the file executable
RUN chmod +x lambda_function.py

# Set the CMD to your handler
CMD [ "lambda_function.lambda_handler" ]</code></pre><p>3. <strong>Build and Push Your Docker Image to Amazon ECR: </strong>Use the AWS CLI to create a repository in ECR, build your Docker image, and push it to the repository. This step requires permissions to interact with ECR and S3.</p><pre><code>aws ecr create-repository - repository-name my-rag-lambda
 docker build -t my-test-lambda .
 docker tag my-rag-lambda:latest &lt;aws_account_id&gt;.dkr.ecr.&lt;region&gt;.amazonaws.com/my-test-lambda:latest
 docker push &lt;aws_account_id&gt;.dkr.ecr.&lt;region&gt;.amazonaws.com/my-test-lambda:latest</code></pre><p>4. <strong>Create a Lambda Function</strong>: In the AWS Lambda console, create a new function using the container image option. Specify the URI of the image in ECR as the image source.</p><p>5. <strong>Configure Your Lambda Function</strong>: Set the necessary environment variables, such as `OPENAI_API_KEY` for LangChain, and configure any other settings as needed.</p><p>6. <strong>Test Your Lambda Function</strong>: Invoke your Lambda function to test it. You can do this from the AWS Lambda console or using the AWS CLI. Ensure that the function is correctly processing inputs and generating outputs as expected.</p><h3>Implementation</h3><p>Below code is designed to deploy a Retrieval-Augmented Generation (RAG) model using AWS Lambda, Docker, and Amazon ECR, with LangChain for language model deployment.</p><h4>Environment Setup</h4><p>First, imports necessary libraries and sets up the environment:</p><pre><code>import boto3
import bs4
from langchain import hub
from langchain_community.document_loaders import WebBaseLoader
from langchain_community.vectorstores import Chroma
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough
from langchain_community.embeddings import OpenAIEmbeddings
from langchain_community.chat_models import ChatOpenAI
from langchain_text_splitters import RecursiveCharacterTextSplitter
import os
# Retrieve the OpenAI API key from environment variables
OPENAI_API_KEY = os.environ['OPENAI_API_KEY']
print(OPENAI_API_KEY)</code></pre><p>- boto3: The AWS SDK for Python, allowing Python developers to write software that makes use of services like Amazon S3, Amazon EC2, and others.<br>- bs4: Beautiful Soup, a library for pulling data out of HTML and XML files.<br>- LangChain: A framework for building and deploying language models, providing tools for document loading, vector storage, embeddings, and more.<br>- Environment Variables: The code retrieves the `OPENAI_API_KEY` from the environment variables, which is crucial for accessing OpenAI&#8217;s API.</p><h4>Data Loading and Processing</h4><p>The `load_data` function is responsible for loading, chunking, and indexing the contents of a web page:</p><pre><code>def load_data():
 loader = WebBaseLoader(
 web_paths=("https://lilianweng.github.io/posts/2023-06-23-agent/",),
 bs_kwargs=dict(
 parse_only=bs4.SoupStrainer(
 class_=("post-content", "post-title", "post-header")
 )
 ),
 )
 docs = loader.load()
 text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
 splits = text_splitter.split_documents(docs)
 vectorstore = Chroma.from_documents(documents=splits, embedding=OpenAIEmbeddings())
 retriever = vectorstore.as_retriever()
 return retriever</code></pre><p>- WebBaseLoader: Loads documents from web paths, using Beautiful Soup to parse and extract specific elements.<br>- RecursiveCharacterTextSplitter: Splits documents into chunks based on character count and overlap.<br>- Chroma: Creates a vector store from the split documents, using OpenAI embeddings for vectorization.<br>- Retriever: A retriever object that can be used to retrieve relevant documents based on queries.</p><h4>Response Generation</h4><p>The `get_response` function generates a response to a given query using the RAG model:</p><pre><code>def get_response(query):
 prompt = hub.pull("rlm/rag-prompt")
 llm = ChatOpenAI(model_name="gpt-3.5-turbo", temperature=0)
 retriever = load_data()
 rag_chain = (
 {"context": retriever | format_docs, "question": RunnablePassthrough()}
 | prompt
 | llm
 | StrOutputParser()
 )
 return rag_chain.invoke(query)</code></pre><p>- hub.pull: Retrieves a prompt from <a href="https://smith.langchain.com/hub/rlm/rag-prompt">LangChain&#8217;s hub</a>.<br>- ChatOpenAI: Initializes a chat model with GPT-3.5 Turbo.<br>- RunnablePassthrough: A component that passes the input directly to the next component in the chain.<br>- StrOutputParser: Parses the output of the RAG model into a string format.</p><h4>AWS Lambda&nbsp;Handler</h4><p>Finally, the `lambda_handler` function is the entry point for AWS Lambda, which receives an event and context, processes the query, and returns a response:</p><pre><code>def lambda_handler(event, context):
 query = event.get("question")
 response = get_response(query)
 print("response:", response)
 return {"body": response, "statusCode": 200}</code></pre><ul><li><p>Event and Context: AWS Lambda passes an event object and a context object to the handler. The event object contains information about the triggering event, and the context object contains information about the runtime environment.</p></li><li><p>Query Processing: The function extracts the query from the event, generates a response using the `get_response` function, and prints the response and returns a response object containing the generated response and a status code of 200, indicating success.</p></li></ul><h3>Conclusion</h3><p>Deploying a RAG model in AWS Lambda using Docker and ECR with LangChain involves preparing your environment, creating a Dockerfile, building and pushing your Docker image to ECR, creating a Lambda function, configuring it, and testing it. This process leverages the serverless capabilities of AWS Lambda, the containerization benefits of Docker, and the language model deployment capability provided by LangChain.</p><blockquote><p><em><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p></blockquote><blockquote><p><em><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></em></p></blockquote><blockquote><p><em><strong>Visit my technical channel on</strong> <strong><a href="https://www.youtube.com/channel/UCGphGM_oeR4r9dqVs71Jc5w">Youtube</a></strong></em></p></blockquote><blockquote><p><em>Thanks for&nbsp;Reading!</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[EDA — Visualize Embeddings of RAG]]></title><description><![CDATA[UMAP &#8212; Visualize RAG data &#8212; Langchain Chroma HuggingFaceEmbeddings]]></description><link>https://aboniasojasingarayar.substack.com/p/eda-visualize-embeddings-of-rag-cf3e0070ccc0</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/eda-visualize-embeddings-of-rag-cf3e0070ccc0</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Wed, 27 Mar 2024 08:32:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ed422dec-0624-44c4-aff1-eefd19b07f1d_1462x720.gif" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>UMAP&#8202;&#8212;&#8202;Visualize RAG data&#8202;&#8212;&#8202;Langchain Chroma HuggingFaceEmbeddings</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!szj0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!szj0!,w_424,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif 424w, https://substackcdn.com/image/fetch/$s_!szj0!,w_848,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif 848w, https://substackcdn.com/image/fetch/$s_!szj0!,w_1272,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif 1272w, https://substackcdn.com/image/fetch/$s_!szj0!,w_1456,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!szj0!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!szj0!,w_424,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif 424w, https://substackcdn.com/image/fetch/$s_!szj0!,w_848,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif 848w, https://substackcdn.com/image/fetch/$s_!szj0!,w_1272,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif 1272w, https://substackcdn.com/image/fetch/$s_!szj0!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96937211-1249-4344-8cae-056444e0e7ba_1462x720.gif 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Courtesy of Spotlight</figcaption></figure></div><p>In this article, we delve into the how we can visualize Retrieval Augmented Generation (RAG) data using the langchain framework in conjunction with Hugging Face&#8217;s language models and embeddings. We&#8217;ll explore how to leverage the `HuggingFaceEmbeddings` and `Chroma` vector store for efficient embedding and document retrieval, followed by a practical example of visualizing these embeddings to understand their distribution and relevance to a given question.</p><p>We&#8217;ll explore how to use the Hugging Face Embeddings for embedding text data and store it in chroma and then visualize it using UMAP (Uniform Manifold Approximation and Projection), a dimensionality reduction technique that helps in visualizing high-dimensional data in a more interpretable way.</p><h4>Introduction to&nbsp;RAG</h4><p>RAG is a cutting-edge approach that combines the strengths of pretrained Large Language Models (LLMs) and your own data to generate responses. It retrieves documents, passes them through a sequence-to-sequence model, and then marginalizes to generate outputs. This method is particularly useful for querying specific documents or interacting with your own data in a conversational manner&nbsp;.</p><h4>Setting Up the Environment</h4><p>To begin, we&#8217;ll need to import the necessary libraries and set up our environment. We&#8217;ll use `pandas` for data manipulation, `langchain.embeddings` for handling embeddings, and `langchain.vectorstores` for our vector store.</p><p><strong><a href="https://python.langchain.com/docs/get_started/introduction" title="https://python.langchain.com/docs/get_started/introduction">Introduction | &#129436;&#65039;&#128279; Langchain</a></strong><a href="https://python.langchain.com/docs/get_started/introduction" title="https://python.langchain.com/docs/get_started/introduction"><br></a><em><a href="https://python.langchain.com/docs/get_started/introduction" title="https://python.langchain.com/docs/get_started/introduction">LangChain is a framework for developing applications powered by language models. It enables applications that:</a></em><a href="https://python.langchain.com/docs/get_started/introduction" title="https://python.langchain.com/docs/get_started/introduction">python.langchain.com</a></p><p>So do install as follow:</p><pre><code>!pip install pandas langchain renumics-spotlight umap-learn</code></pre><pre><code>import pandas as pd
from langchain.embeddings import HuggingFaceEmbeddings
from langchain.vectorstores import Chroma</code></pre><h4>Loading the Model and Vector&nbsp;Store</h4><p>Next, we&#8217;ll specify the model path for our embeddings and create instances of `HuggingFaceEmbeddings` and `Chroma`. We&#8217;ll also configure the model to use the CPU/GPU for computations and ensure embeddings are not normalized for our visualization purposes.</p><p>Before dive in we consider, you already have chroma db collection and see how we can visualize it.If no you can create one as follow:</p><pre><code>from langchain_community.document_loaders.recursive_url_loader import RecursiveUrlLoader
from bs4 import BeautifulSoup as Soup
from langchain.vectorstores import Chroma

url = "https://www.freenews.fr/"
loader = RecursiveUrlLoader(
    url=url, max_depth=5, extractor=lambda x: Soup(x, "lxml").text
)
documents = loader.load()
text_splitter = RecursiveCharacterTextSplitter(
            chunk_size=1000, chunk_overlap=0
        )
texts = text_splitter.split_documents(documents)

modelPath = "sentence-transformers/distiluse-base-multilingual-cased-v1"

model_kwargs = {'device':'cpu'}
#encode_kwargs = {'normalize_embeddings': False}

embeddings = HuggingFaceEmbeddings(model_name=modelPath,model_kwargs=model_kwargs) 

# SAVE
docs_vectorstore = Chroma.from_documents(texts, embeddings, persist_directory="./chroma_db_multilingual")</code></pre><p>Provide the link to your chroma persist_directory below:</p><pre><code>modelPath = "sentence-transformers/distiluse-base-multilingual-cased-v1"
model_kwargs = {'device':'cpu'}
encode_kwargs = {'normalize_embeddings': False}
embeddings_model = HuggingFaceEmbeddings(model_name=modelPath, model_kwargs=model_kwargs)
docs_vectorstore = Chroma(persist_directory="./chroma_db_multilingual", embedding_function=embeddings_model)</code></pre><h4>Fetching and Preparing Data</h4><p>We&#8217;ll retrieve our data from the vector store, including metadata, documents, and embeddings. Then, we&#8217;ll create a DataFrame to organize this data and add a column to indicate whether each document contains the answer to our query.</p><pre><code>response = docs_vectorstore.get(include=["metadatas", "documents", "embeddings"])
df = pd.DataFrame({
 "id": response["ids"],
 "source": [metadata.get("source") for metadata in response["metadatas"]],
 "page": [metadata.get("page", -1) for metadata in response["metadatas"]],
 "document": response["documents"],
 "embedding": response["embeddings"],
})
df["contains_answer"] = df["document"].apply(lambda x: "N&#339;ud R&#233;partition Optique)" in x)</code></pre><h3>Calculating Distances</h3><p>To find the closest match to our question, we calculate the Euclidean distance between the question embedding and each document embedding.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!taJ7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!taJ7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png 424w, https://substackcdn.com/image/fetch/$s_!taJ7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png 848w, https://substackcdn.com/image/fetch/$s_!taJ7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png 1272w, https://substackcdn.com/image/fetch/$s_!taJ7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!taJ7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!taJ7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png 424w, https://substackcdn.com/image/fetch/$s_!taJ7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png 848w, https://substackcdn.com/image/fetch/$s_!taJ7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png 1272w, https://substackcdn.com/image/fetch/$s_!taJ7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbead4f7c-3992-4620-b3fd-24a0d3904aa3_800x577.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Euclidean Distance</figcaption></figure></div><pre><code>question_embedding = embeddings_model.embed_query(question)
df["dist"] = df.apply(
    lambda row: np.linalg.norm(
        np.array(row["embedding"]) - question_embedding
    ),
    axis=1,
)</code></pre><h3>Visualizing Embeddings</h3><h4>1. Using Spotlight</h4><p>This visualization will help us understand how closely related documents are to our query and identify potential areas of improvement or further investigation.</p><p>Finally, we use the <code>spotlight</code> module from <code>renumics</code> to visualize the data. This step is crucial for understanding the distribution of documents in relation to the question and the model's response.</p><p><strong><a href="https://github.com/Renumics/spotlight" title="https://github.com/Renumics/spotlight">GitHub - Renumics/spotlight: Interactively explore unstructured datasets from your dataframe.</a></strong><a href="https://github.com/Renumics/spotlight" title="https://github.com/Renumics/spotlight"><br></a><em><a href="https://github.com/Renumics/spotlight" title="https://github.com/Renumics/spotlight">Interactively explore unstructured datasets from your dataframe. - Renumics/spotlight</a></em><a href="https://github.com/Renumics/spotlight" title="https://github.com/Renumics/spotlight">github.com</a></p><pre><code>from renumics import spotlight
spotlight.show(df)</code></pre><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mEIP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mEIP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png 424w, https://substackcdn.com/image/fetch/$s_!mEIP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png 848w, https://substackcdn.com/image/fetch/$s_!mEIP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png 1272w, https://substackcdn.com/image/fetch/$s_!mEIP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mEIP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mEIP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png 424w, https://substackcdn.com/image/fetch/$s_!mEIP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png 848w, https://substackcdn.com/image/fetch/$s_!mEIP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png 1272w, https://substackcdn.com/image/fetch/$s_!mEIP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa6fb367f-8b00-421a-84ed-61fe6492591a_800x369.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">UMAP&#8202;&#8212;&#8202;Visualize the embeddings</figcaption></figure></div><p><strong><a href="https://github.com/Renumics/rag-demo/blob/main/notebooks/visualize_rag_tutorial.ipynb" title="https://github.com/Renumics/rag-demo/blob/main/notebooks/visualize_rag_tutorial.ipynb">rag-demo/notebooks/visualize_rag_tutorial.ipynb at main &#183; Renumics/rag-demo</a></strong><a href="https://github.com/Renumics/rag-demo/blob/main/notebooks/visualize_rag_tutorial.ipynb" title="https://github.com/Renumics/rag-demo/blob/main/notebooks/visualize_rag_tutorial.ipynb"><br></a><em><a href="https://github.com/Renumics/rag-demo/blob/main/notebooks/visualize_rag_tutorial.ipynb" title="https://github.com/Renumics/rag-demo/blob/main/notebooks/visualize_rag_tutorial.ipynb">Retrieval-Augmented Generation Assistant Demo &#129302;&#10133;&#128218;&#129008;&#10084;&#65039; - rag-demo/notebooks/visualize_rag_tutorial.ipynb at main &#183;&#8230;</a></em><a href="https://github.com/Renumics/rag-demo/blob/main/notebooks/visualize_rag_tutorial.ipynb" title="https://github.com/Renumics/rag-demo/blob/main/notebooks/visualize_rag_tutorial.ipynb">github.com</a></p><h4>2. Using&nbsp;UMAP</h4><pre><code>import umap
# Find the  5 closest vectors
closest_vectors_indices = df.nsmallest(5, 'dist')['id'].values

# Prepare the embeddings for UMAP
embeddings = np.array([np.array(x) for x in df["embedding"]])

# Reduce dimensionality with UMAP
reducer = umap.UMAP()
embedding_reduced = reducer.fit_transform(embeddings)

# Plot the reduced embeddings
plt.scatter(embedding_reduced[:,  0], embedding_reduced[:,  1], c='gray', alpha=0.2)

# Highlight the question embedding and the  5 closest vectors
plt.scatter(embedding_reduced[df["id"].isin(closest_vectors_indices),  0], embedding_reduced[df["id"].isin(closest_vectors_indices),  1], c='red', alpha=1)
plt.scatter(embedding_reduced[df["id"] == df[df["dist"] == df["dist"].min()]["id"].values[0],  0], embedding_reduced[df["id"] == df[df["dist"] == df["dist"].min()]["id"].values[0],  1], c='blue', alpha=1, marker='*')

# Add labels and title
plt.title("UMAP Visualization of Text Embeddings with Question Highlighted")
plt.xlabel("UMAP  1")
plt.ylabel("UMAP  2")
plt.show()</code></pre><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Z58H!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Z58H!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png 424w, https://substackcdn.com/image/fetch/$s_!Z58H!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png 848w, https://substackcdn.com/image/fetch/$s_!Z58H!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png 1272w, https://substackcdn.com/image/fetch/$s_!Z58H!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Z58H!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Z58H!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png 424w, https://substackcdn.com/image/fetch/$s_!Z58H!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png 848w, https://substackcdn.com/image/fetch/$s_!Z58H!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png 1272w, https://substackcdn.com/image/fetch/$s_!Z58H!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e2c085f-34f4-4d0a-a4d8-ae00b1fb0e64_588x455.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">UMAP&#8202;&#8212;&#8202;Visualize Text embedding&#8202;&#8212;&#8202;Blue point is the&nbsp;Question</figcaption></figure></div><h4>3. Using Tensorboard</h4><p>Another effective way to visualize embeddings, especially useful for gaining insights into word embeddings and the relationships between them, is by using TensorBoard. TensorBoard is a tool that allows for the visualization of machine learning models and their metrics, including word embeddings, in a user-friendly interface. It can be particularly beneficial when you want to explore the semantic similarities and relationships between words in your embeddings.</p><p><strong>Setting Up TensorBoard</strong></p><p>First, ensure you have TensorBoard installed. You can install it using pip:</p><pre><code>pip install tensorboard</code></pre><p><strong>Visualizing Word Embeddings with TensorBoard</strong></p><p>To visualize word embeddings using TensorBoard, you&#8217;ll typically need to convert your embeddings into a format that TensorBoard can understand. This often involves creating a metadata file that maps words to their embeddings and then using TensorBoard&#8217;s `Projector` to visualize these embeddings.</p><p><strong><a href="https://www.tensorflow.org/tensorboard/get_started" title="https://www.tensorflow.org/tensorboard/get_started">Get started with TensorBoard | TensorFlow</a></strong><a href="https://www.tensorflow.org/tensorboard/get_started" title="https://www.tensorflow.org/tensorboard/get_started"><br></a><em><a href="https://www.tensorflow.org/tensorboard/get_started" title="https://www.tensorflow.org/tensorboard/get_started">In machine learning, to improve something you often need to be able to measure it. TensorBoard is a tool for providing&#8230;</a></em><a href="https://www.tensorflow.org/tensorboard/get_started" title="https://www.tensorflow.org/tensorboard/get_started">www.tensorflow.org</a></p><p>Here&#8217;s a simplified step-by-step guide on how to do this:</p><p>1. Prepare Your Embeddings: Ensure your embeddings are in a suitable format. For TensorBoard, you might need to convert your embeddings into a `.tsv` (Tab-Separated Values) file where each line contains a word and its corresponding embedding vector.</p><p>2. Create a Metadata File: Alongside your embeddings file, create a metadata file that maps each word to its index in the embeddings file. This file is also in `.tsv` format.</p><p>3. Use TensorBoard&#8217;s Projector:<br>&#8202;&#8212;&#8202;Start TensorBoard by running `tensorboard&#8202;&#8212;&#8202;logdir=path/to/your/logs` in your terminal.<br>&#8202;&#8212;&#8202;In your web browser, navigate to the TensorBoard interface (usually at `localhost:6006`).<br>&#8202;&#8212;&#8202;Go to the `Projector` tab.<br>&#8202;&#8212;&#8202;Click on `Load` under the `Embeddings` section.<br>&#8202;&#8212;&#8202;Upload your embeddings file and metadata file.</p><p>4. Explore Your Embeddings:<br>&#8202;&#8212;&#8202;Once your embeddings are loaded, you can explore them in various ways, such as through a 2D or 3D scatter plot, where each point represents a word and its position is determined by its embedding vector.<br>&#8202;&#8212;&#8202;You can also use the `Word` search bar to find specific words and see how they are positioned relative to others in the embedding space.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YouH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YouH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png 424w, https://substackcdn.com/image/fetch/$s_!YouH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png 848w, https://substackcdn.com/image/fetch/$s_!YouH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png 1272w, https://substackcdn.com/image/fetch/$s_!YouH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YouH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YouH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png 424w, https://substackcdn.com/image/fetch/$s_!YouH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png 848w, https://substackcdn.com/image/fetch/$s_!YouH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png 1272w, https://substackcdn.com/image/fetch/$s_!YouH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44cda0e8-83d4-49cf-8802-baa9c0ec993a_800x380.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>-Semantic Similarity: TensorBoard makes it easier to visually inspect the semantic relationships between words by examining their positions in the embedding space.<br>- Ease of Use: The user-friendly interface of TensorBoard allows for intuitive exploration of embeddings without the need for complex code.<br>- Insight into Model Performance: Beyond visualizing embeddings, TensorBoard can also be used to track model metrics, network weight distributions, and other performance indicators, providing a comprehensive toolkit for monitoring and understanding your models.</p><h4>Conclusion</h4><p>Visualizing RAG data provides valuable insights into the model&#8217;s decision-making process, helping us understand how it selects and interprets information from external documents.By visualizing RAG data, we gain valuable insights into the relationships between documents and our queries. This technique not only aids in understanding the performance of our RAG system but also guides us in refining models and datasets for better accuracy and relevance.</p><blockquote><p><em><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p></blockquote><blockquote><p><em><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></em></p></blockquote><blockquote><p><strong>Visit my technical channel on</strong> <strong><a href="http://www.youtube.com/@aboniasojasingarayar3097">Youtube</a></strong></p></blockquote><blockquote><p>Thanks for&nbsp;Reading!</p></blockquote>]]></content:encoded></item><item><title><![CDATA[Build and Deploy LLM Application in AWS]]></title><description><![CDATA[AWS Lambda &#8212; BedRock &#8212; LangChain]]></description><link>https://aboniasojasingarayar.substack.com/p/build-and-deploy-llm-application-in-aws-cca46c662749</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/build-and-deploy-llm-application-in-aws-cca46c662749</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Mon, 18 Mar 2024 08:36:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/HiEjhVc_Dzc" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>AWS Lambda&#8202;&#8212;&#8202;BedRock&#8202;&#8212;&#8202;LangChain</h4><div class="captioned-image-container"><figure><div id="youtube2-HiEjhVc_Dzc" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;HiEjhVc_Dzc&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/HiEjhVc_Dzc?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><figcaption class="image-caption">Tutorial by Author</figcaption></figure></div><p>This article helps to explore the resource and implementation details that we covered in above tutorial, we&#8217;ll explore how to build and deploy a Large Language Model (LLM) application on AWS Lambda, leveraging the power of BedRock and LangChain. This application will allow users to ask natural language questions.</p><p>We collaborated with <strong><a href="https://chetanhirapara.medium.com/">Chetan Hirapara</a></strong> on this tutorial. Feel free to check out his<strong> <a href="https://www.linkedin.com/in/chetan-hirapara-90344345/?originalSubdomain=in">Linkedin</a></strong> profile.</p><p>We have covered the following topic in the tutorial:</p><h4>Introduction</h4><p>AWS Lambda is a serverless computing service that runs your code in response to events and automatically manages the underlying compute resources for you. BedRock is a service that provides serverless access to foundational models, including LLMs, while LangChain is a framework that integrates with these models to facilitate conversational interactions.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GR4I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GR4I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!GR4I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!GR4I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!GR4I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GR4I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GR4I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png 424w, https://substackcdn.com/image/fetch/$s_!GR4I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png 848w, https://substackcdn.com/image/fetch/$s_!GR4I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png 1272w, https://substackcdn.com/image/fetch/$s_!GR4I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ac02d7f-8a72-424e-adbd-ac1d0f13b6bd_800x450.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Image by&nbsp;Author</figcaption></figure></div><h4>Bedrock and Lambda&nbsp;Service</h4><p>To start, we&#8217;ll set up our AWS Lambda function to interact with BedRock. BedRock allows us to access a variety of LLMs, including those developed by leading AI startups like AI21 Labs, Anthropic, and Cohere. This setup is crucial for our application, as it enables us to leverage the text generation and analysis capabilities of these models.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QBVP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QBVP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg 424w, https://substackcdn.com/image/fetch/$s_!QBVP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg 848w, https://substackcdn.com/image/fetch/$s_!QBVP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!QBVP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QBVP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QBVP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg 424w, https://substackcdn.com/image/fetch/$s_!QBVP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg 848w, https://substackcdn.com/image/fetch/$s_!QBVP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!QBVP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43d3888b-a5b1-4fd5-aa77-ad1c4c2f6c2c_800x379.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Architecture</figcaption></figure></div><h4>Create Lambda&nbsp;Function</h4><p>Next, we&#8217;ll create a Lambda function that will handle the processing of user queries.</p><p><strong>Code and Implementation</strong></p><pre><code>import boto3
from langchain_community.embeddings import BedrockEmbeddings
from langchain.llms.bedrock import Bedrock
from langchain.prompts import PromptTemplate
from langchain.chains import LLMChain

# Bedrock Client Setup
bedrock = boto3.client(service_name="bedrock-runtime", region_name="us-east-1")
def get_llama2_llm():
    llm = Bedrock(
        model_id="meta.llama2-13b-chat-v1",
        client=bedrock,
        model_kwargs={"max_gen_len": 512},
    )
    return llm
    
def get_response_llm(llm, query):
    prompt_template = """&lt;s&gt;[INST] You are a helpful, respectful and honest assistant.
    {question} [/INST] &lt;/s&gt;
    """
    prompt = PromptTemplate(
        template=prompt_template, input_variables=["question"]
    )
    chain = LLMChain(llm=llm, prompt=prompt)
    return chain.invoke(query)
# Lambda Handler
def lambda_handler(event, context):
    user_question = event.get("question")
    llm = get_llama2_llm()
    response = get_response_llm(llm, user_question)
    print("response: ", response)
    return {"body": response, "statusCode": 200}</code></pre><h4>Create and Add Custom Lambda&nbsp;Layer</h4><p>To enhance our application, we might need to add custom Lambda layers. These layers can include additional libraries or dependencies that our Lambda function requires. For example, we might need to include the `boto3` library for interacting with AWS services or the `langchain` library for integrating with LangChain.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3GvK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3GvK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png 424w, https://substackcdn.com/image/fetch/$s_!3GvK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png 848w, https://substackcdn.com/image/fetch/$s_!3GvK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png 1272w, https://substackcdn.com/image/fetch/$s_!3GvK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3GvK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3GvK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png 424w, https://substackcdn.com/image/fetch/$s_!3GvK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png 848w, https://substackcdn.com/image/fetch/$s_!3GvK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png 1272w, https://substackcdn.com/image/fetch/$s_!3GvK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F344da894-f312-4f0f-8dab-ee87637abf5a_800x455.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Lambda Layer&#8202;&#8212;&#8202;Image by&nbsp;Author</figcaption></figure></div><h4>Test</h4><p>Once the Lambda function is deployed with the necessary layers, we&#8217;ll test it again to ensure it&#8217;s correctly processing user queries and generating responses. This involves sending prompts to the backend and verifying that the LLM generates appropriate responses.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LoxK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LoxK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png 424w, https://substackcdn.com/image/fetch/$s_!LoxK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png 848w, https://substackcdn.com/image/fetch/$s_!LoxK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png 1272w, https://substackcdn.com/image/fetch/$s_!LoxK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LoxK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LoxK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png 424w, https://substackcdn.com/image/fetch/$s_!LoxK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png 848w, https://substackcdn.com/image/fetch/$s_!LoxK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png 1272w, https://substackcdn.com/image/fetch/$s_!LoxK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71af5813-9628-4f73-8ac4-ea2de8a22ad6_800x455.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Deployment and Test&#8202;&#8212;&#8202;Image by&nbsp;Author</figcaption></figure></div><p>In summary, building and deploying a serverless LLM application on AWS Lambda using BedRock and LangChain involves setting up a Lambda function, deploying it, and testing it to ensure it can process user queries and generate responses. This solution combines the capabilities of LLMs and semantic search to answer natural language questions&nbsp;, serving as a blueprint for further generative AI use cases.</p><h4>&#128218; Resources</h4><blockquote><p>- <strong>ARN for popular packages:</strong> <a href="https://api.klayers.cloud/api/v2/p3.11/layers/latest/us-east-1/html]%28https://api.klayers.cloud/api/v2/p3.11/layers/latest/us-east-1/html%29">https://api.klayers.cloud/api/v2/p3.11/layers/latest/us-east-1/html</a><br>- <strong>Amazon ECR Public Gallery&#8202;&#8212;&#8202;SAM: </strong><a href="https://gallery.ecr.aws/sam/build-python3.10]%28https://gallery.ecr.aws/sam/build-python3.10%29">https://gallery.ecr.aws/sam/build-python3.10</a><br>- <strong>Official AWS Documentation</strong>: <a href="https://aws.amazon.com/blogs/compute/building-a-serverless-document-chat-with-aws-lambda-and-amazon-bedrock/]%28https://aws.amazon.com/blogs/compute/building-a-serverless-document-chat-with-aws-lambda-and-amazon-bedrock/%29">https://aws.amazon.com/blogs/compute/building-a-serverless-document-chat-with-aws-lambda-and-amazon-bedrock/</a><br>- <strong>Docker</strong>: <a href="https://docs.docker.com/desktop/install/mac-install/]%28https://docs.docker.com/desktop/install/mac-install/%29">https://docs.docker.com/desktop/install/mac-install/</a></p></blockquote><h3>Conclusion</h3><p>This tutorial provides a comprehensive guide to building and deploying a serverless LLM application on AWS Lambda, leveraging the power of BedRock and LangChain. By following these steps, you&#8217;ll be able to create an application that allows users to ask natural language questions, showcasing the potential of serverless computing and generative AI.</p><div class="captioned-image-container"><figure><div id="youtube2-HiEjhVc_Dzc" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;HiEjhVc_Dzc&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/HiEjhVc_Dzc?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><figcaption class="image-caption">Tutorial by Author</figcaption></figure></div><blockquote><p><em><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p></blockquote><blockquote><p><em><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></em></p></blockquote><blockquote><p><em><strong>Visit my technical channel on</strong> <strong><a href="https://www.youtube.com/channel/UCGphGM_oeR4r9dqVs71Jc5w">Youtube</a></strong></em></p></blockquote><blockquote><p><em>Thanks for&nbsp;Reading!</em></p></blockquote>]]></content:encoded></item><item><title><![CDATA[Ollama and LangChain: Run LLMs locally]]></title><description><![CDATA[Run open-source LLM, such as Llama 2,mistral locally]]></description><link>https://aboniasojasingarayar.substack.com/p/ollama-and-langchain-run-llms-locally-900931914a46</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/ollama-and-langchain-run-llms-locally-900931914a46</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Thu, 29 Feb 2024 08:28:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/hO1cGqXGM8k" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Run open-source LLM, such as Llama 2,mistral locally</h4><div class="captioned-image-container"><figure><div id="youtube2-hO1cGqXGM8k" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;hO1cGqXGM8k&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/hO1cGqXGM8k?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><figcaption class="image-caption">Ollma+Langchain&#8202;&#8212;&#8202;Demo</figcaption></figure></div><h3>Introduction</h3><p>In the realm of Large Language Models (LLMs), Ollama and LangChain emerge as powerful tools for developers and researchers. Ollama provides a seamless way to run open-source LLMs locally, while LangChain offers a flexible framework for integrating these models into applications. This article will guide you through the process of setting up and utilizing Ollama and LangChain, enabling you to harness the power of LLMs for your projects.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ozr9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ozr9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin 424w, https://substackcdn.com/image/fetch/$s_!ozr9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin 848w, https://substackcdn.com/image/fetch/$s_!ozr9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin 1272w, https://substackcdn.com/image/fetch/$s_!ozr9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ozr9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ozr9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin 424w, https://substackcdn.com/image/fetch/$s_!ozr9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin 848w, https://substackcdn.com/image/fetch/$s_!ozr9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin 1272w, https://substackcdn.com/image/fetch/$s_!ozr9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ba5f21e-009d-4ba7-8835-a29745ca15f6_800x600.bin 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h3>1. Setting Up&nbsp;Ollama</h3><h4>Installation and Configuration</h4><p>To start using Ollama, you first need to install it on your system. For macOS users, Homebrew simplifies this process:</p><pre><code>brew install ollama
brew services start ollama</code></pre><p>After installation, Ollama listens on port 11434 for incoming requests. You can verify its operation by navigating to `<a href="http://localhost:11434/`">http://localhost:11434/`</a> in your browser. The next step involves selecting the LLM you wish to run locally. For instance, to run the `llama2`, `mistral` etc model, execute:</p><pre><code>ollama pull mistral</code></pre><p>This command downloads the model, optimizing setup and configuration details, including GPU usage.</p><p>Also you can download and install ollama from <a href="https://ollama.com/download/linux">official site</a>.</p><h3>2. Running&nbsp;Models</h3><p>To interact with your locally hosted LLM, you can use the command line directly or via an API. For command-line interaction, Ollama provides the `ollama run &lt;name-of-model&gt;` command. Alternatively, you can send a JSON request to the API endpoint of Ollama:</p><pre><code>curl http://localhost:11434/api/generate -d '{
 "model": "llama2",
 "prompt":"Why is the sky blue?"
}'</code></pre><p>This flexibility allows you to integrate LLMs into various applications seamlessly.</p><h3>3. Integrating Ollama with LangChain</h3><p>LangChain is a framework designed to facilitate the integration of LLMs into applications. It supports a wide range of chat models, including Ollama, and provides an expressive language for chaining operations. To get started, you&#8217;ll need to install LangChain and its dependencies.</p><blockquote><p>Official Documenation</p></blockquote><p><strong><a href="https://python.langchain.com/docs/integrations/llms/ollama" title="https://python.langchain.com/docs/integrations/llms/ollama">Ollama | &#129436;&#65039;&#128279; Langchain</a></strong><a href="https://python.langchain.com/docs/integrations/llms/ollama" title="https://python.langchain.com/docs/integrations/llms/ollama"><br></a><em><a href="https://python.langchain.com/docs/integrations/llms/ollama" title="https://python.langchain.com/docs/integrations/llms/ollama">Ollama allows you to run open-source large</a></em><a href="https://python.langchain.com/docs/integrations/llms/ollama" title="https://python.langchain.com/docs/integrations/llms/ollama">python.langchain.com</a></p><h4>Using Ollama in LangChain</h4><p>To use Ollama within a LangChain application, you first import the necessary modules from the `langchain_community.llms` package:</p><pre><code>from langchain_community.llms import Ollama</code></pre><p>Then, initialize an instance of the Ollama model:</p><pre><code>llm = Ollama(model="llama2")</code></pre><p>You can now invoke the model to generate responses. For example:</p><pre><code>llm.invoke("Tell me a joke")</code></pre><p>This code snippet demonstrates how to use Ollama to generate a response to a given prompt.</p><h4>Advanced Usage</h4><pre><code>from langchain_community.llms import Ollama

llm = Ollama(model="mistral")
llm("The first man on the summit of Mount Everest, the highest peak on Earth, was ...")</code></pre><p>LangChain also supports more complex operations, such as streaming responses and using prompt templates. For instance, you can stream responses from the model as follows:</p><pre><code>from langchain.callbacks.manager import CallbackManager
from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler

llm = Ollama(
    model="mistral", callback_manager=CallbackManager([StreamingStdOutCallbackHandler()])
)
llm("The first man on the summit of Mount Everest, the highest peak on Earth, was ...")</code></pre><p>This approach is particularly useful for applications requiring real-time interaction with LLMs.</p><h3>4. Deploying with LangServe</h3><p>For production environments, LangChain offers LangServe, a deployment tool that simplifies the process of running your application. You can use LangServe to deploy your LangChain application.</p><p><strong><a href="https://github.com/langchain-ai/langserve">LangServe</a> is an open-source library of LangChain</strong> that makes your process for creating API servers based on your chains easier. LangServe provides remote APIs for core LangChain Expression Language methods such as invoke, batch, and stream.</p><pre><code>from typing import List
from fastapi import FastAPI
from langchain.llms import Ollama
from langchain.output_parsers import CommaSeparatedListOutputParser
from langchain.prompts import PromptTemplate
from langserve import add_routes
import uvicorn

llama2 = Ollama(model="mistral")
template = PromptTemplate.from_template("Tell me a joke about {topic}.")
chain = template | llama2 | CommaSeparatedListOutputParser()

app = FastAPI(title="LangChain", version="1.0", description="The first server ever!")
add_routes(app, chain, path="/chain")

if __name__ == "__main__":
    uvicorn.run(app, host="localhost", port=8000)</code></pre><p>Run the below code to start your lagserve and head on to</p><blockquote><p><em>http://localhost:9001/chain/playground/</em></p></blockquote><p>to access your playground for to generate the joke on particular topic.Its your turn now to test with your custom prompt.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2vIO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2vIO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png 424w, https://substackcdn.com/image/fetch/$s_!2vIO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png 848w, https://substackcdn.com/image/fetch/$s_!2vIO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png 1272w, https://substackcdn.com/image/fetch/$s_!2vIO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2vIO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2vIO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png 424w, https://substackcdn.com/image/fetch/$s_!2vIO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png 848w, https://substackcdn.com/image/fetch/$s_!2vIO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png 1272w, https://substackcdn.com/image/fetch/$s_!2vIO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4d7ea9d9-3754-469c-8bc3-8ac77bf01e34_1200x334.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Launch LangServe</figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!v04I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!v04I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png 424w, https://substackcdn.com/image/fetch/$s_!v04I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png 848w, https://substackcdn.com/image/fetch/$s_!v04I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png 1272w, https://substackcdn.com/image/fetch/$s_!v04I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!v04I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!v04I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png 424w, https://substackcdn.com/image/fetch/$s_!v04I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png 848w, https://substackcdn.com/image/fetch/$s_!v04I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png 1272w, https://substackcdn.com/image/fetch/$s_!v04I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb3861c1d-f29f-4a20-8aee-879ad41da546_800x606.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Demo-Playground</figcaption></figure></div><h3>Conclusion</h3><p>By integrating Ollama with LangChain, developers can leverage the capabilities of LLMs without the need for external APIs. This setup not only saves costs but also allows for greater flexibility and customization. Whether you&#8217;re building a chatbot, a content generation tool, or an interactive application, Ollama and LangChain provide the tools necessary to bring LLMs to life.</p><blockquote><p><em><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p></blockquote><blockquote><p><em><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></em></p></blockquote><blockquote><p><strong>Visit my technical channel on</strong> <strong><a href="https://www.youtube.com/channel/UCGphGM_oeR4r9dqVs71Jc5w">Youtube</a></strong></p></blockquote><blockquote><p>Thanks for&nbsp;Reading!</p></blockquote>]]></content:encoded></item><item><title><![CDATA[Machine Learning Model Serving Framework]]></title><description><![CDATA[Framework/Server for model serving]]></description><link>https://aboniasojasingarayar.substack.com/p/machine-learning-model-serving-framework-13945633ecf3</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/machine-learning-model-serving-framework-13945633ecf3</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Wed, 11 Oct 2023 07:46:21 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ba19566d-205b-4ffb-8cb0-626e985a3fd5_800x600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Framework/Server for model&nbsp;serving</h4><p>Machine learning models or LLM have become an integral part of many industries. As these models become more prevalent, the need to serve these models in production environments has become increasingly important. Serving a machine learning model means making it available for prediction or inference, either to serve real-time predictions to users or to batch prediction for offline use. This article will provide an overview of various frameworks and servers used for serving machine learning models and their trade-offs.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FRuH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FRuH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png 424w, https://substackcdn.com/image/fetch/$s_!FRuH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png 848w, https://substackcdn.com/image/fetch/$s_!FRuH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png 1272w, https://substackcdn.com/image/fetch/$s_!FRuH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FRuH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FRuH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png 424w, https://substackcdn.com/image/fetch/$s_!FRuH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png 848w, https://substackcdn.com/image/fetch/$s_!FRuH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png 1272w, https://substackcdn.com/image/fetch/$s_!FRuH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7654942f-d084-4e3d-b146-fb156a2ddcf3_800x600.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Image by&nbsp;Author</figcaption></figure></div><h3>Frameworks for Model&nbsp;Serving</h3><ol><li><p><a href="https://github.com/bentoml/BentoML">BentoML</a>: A framework for building reliable, scalable, and cost-efficient AI applications. It comes with everything you need for model serving, application packaging, and production deployment.</p></li><li><p><a href="https://jina.ai/">Jina</a>: Build multimodal AI services via cloud native technologies. It provides features like Model Serving, Generative AI, Neural Search, and Cloud Native.</p></li><li><p><a href="https://github.com/mosecorg/mosec">Mosec</a>: A machine learning model serving framework with dynamic batching and pipelined stages, providing an easy-to-use Python interface.</p></li><li><p><a href="https://www.tensorflow.org/tfx/guide/serving">TFServing</a>: A flexible, high-performance serving system for machine learning models.</p></li><li><p><a href="https://pytorch.org/serve/">Torchserve</a>: Serve, optimize and scale PyTorch models in production.</p></li><li><p><a href="https://github.com/triton-inference-server/server">Triton Server (TRTIS)</a>: The Triton Inference Server provides an optimized cloud and edge inferencing solution.</p></li><li><p><a href="https://github.com/cortexlabs/cortex">Cortex</a>: An open-source platform for deploying, managing, and scaling machine learning models. It supports deployment of all types of models.</p></li><li><p><a href="https://github.com/kubeflow/kfserving">KFServing</a>: Provides a Kubernetes Custom Resource Definition (CRD) for serving machine learning models on arbitrary frameworks.</p></li><li><p><a href="https://github.com/aws/multi-model-server">Multi Model Server</a>: A flexible and easy-to-use tool for serving deep learning models trained using any ML/DL framework.</p></li><li><p><a href="https://github.com/xinference/xinference">Xinference</a>: Replace OpenAI GPT with another LLM in your app by changing a single line of code. Xinference gives you the freedom to use any LLM you need.</p></li><li><p><a href="https://github.com/lanarky/lanarky">Lanarky</a>: FastAPI framework to build production-grade LLM applications&nbsp;.</p></li><li><p><a href="https://github.com/langchain/langchain-serve">Langchain-serve</a>: Serverless LLM apps on Production with Jina AI Cloud.</p></li><li><p><a href="https://github.com/SeldonIO/seldon-core">Seldon Core</a>: An open-source platform for deploying, scaling, and managing machine learning models in Kubernetes.</p></li><li><p><a href="https://docs.ray.io/en/latest/serve/index.html">Ray Serve</a>: A scalable and programmable serving framework built on top of Ray to help you scale your microservices and ML models in production.</p></li><li><p><a href="https://opensearch.org/docs/2.4/ml-commons-plugin/model-serving-framework/">OpenSearch ML Commons</a>: An experimental feature that allows you to serve custom models and use those models to make inferences.</p></li><li><p><a href="https://kserve.github.io/website/0.10/modelserving/data_plane/data_plane/">KServe</a>: A Kubernetes-native platform to deploy and serve machine learning models&nbsp;.</p></li><li><p><a href="https://www.mlflow.org/docs/latest/models.html">MLflow Model Serving</a>: A flexible, high-performance serving layer for machine learning models built using PyFunc&nbsp;.</p></li><li><p><a href="https://www.kubeflow.org/">Kubeflow</a>: A machine learning toolkit for Kubernetes which aims to make deployments of machine learning workflows on Kubernetes simple, portable, and scalable&nbsp;.</p></li><li><p><a href="https://scikit-learn.org/stable/modules/model_persistence.html">Scikit-learn</a>: A machine learning library for Python which supports model persistence, allowing you to save and load models&nbsp;.</p></li><li><p><a href="https://www.h2o.ai/">H2O</a>: An open-source platform for data analysis, machine learning, and predictive modeling&nbsp;.</p></li><li><p><a href="https://fastapi.tiangolo.com/">FastAPI</a>: A modern, fast (high-performance), web framework for building APIs with Python 3.6+ based on standard Python type hints&nbsp;.</p></li><li><p><a href="https://flask.palletsprojects.com/en/2.0.x/">Flask</a>: A lightweight WSGI web application framework. It is designed to make getting started quick and easy, with the ability to scale up to complex applications&nbsp;.</p></li><li><p><a href="https://pypi.org/project/mlserver/">MLServer</a>: An open-source inference server for machine learning models. It provides a unified API for serving models and supports multiple inference runtimes&nbsp;.</p></li><li><p><a href="https://neptune.ai/">Neptune</a>: A platform for tracking, comparing, and sharing machine learning metadata. It can be used to manage the machine learning lifecycle, including tracking experiments, reproducibility, and deployment.</p></li></ol><p>In conclusion, the field of machine learning is vast and continually evolving. The tools and technologies available for serving machine learning models are diverse and continually improving. While each tool has its strengths and weaknesses, all play a crucial role in enabling the practical application of machine learning models in real-world scenarios. As developers, understanding these tools and their capabilities is key to effectively utilizing machine learning in production environments.</p><h3>Futher References</h3><p><a href="https://ckaestne.medium.com/machine-learning-in-production-from-models-to-systems-e1422ec7cd65">ckaestne.medium.com</a></p><p><a href="https://www.anyscale.com/blog/serving-ml-models-in-production-common-patterns">MLOPS-anyscale.com</a></p><p><a href="https://nanonets.com/blog/machine-learning-production-retraining/">nanonets.com</a></p><p><a href="https://medium.com/@behnaz.nojavan/top-10-tools-for-deploying-machine-learning-models-to-production-pros-cons-and-features-244daf125f02">behnaz.medium.com</a></p><p><a href="https://www.sigmoid.com/blogs/5-best-practices-for-putting-ml-models-into-production/">sigmoid.com</a></p><p><a href="https://www.ncbi.nlm.nih.gov/books/NBK109714/">ncbi.nlm.nih.gov</a></p><p><a href="https://blog.devgenius.io/deploying-a-machine-learning-model-using-tensorflow-serving-and-docker-3a58f774b30e">blog.devgenius.io</a></p><p><a href="https://llego.dev/posts/production-machine-learning/">llego.dev</a></p><p><a href="https://censius.ai/blogs/things-to-consider-for-model-serving">censius.ai</a></p><p><a href="https://saturncloud.io/blog/exporting-machine-learning-models-a-comprehensive-guide-for-data-scientists/">saturncloud.io</a></p><blockquote><p><em><strong>Connect with me on <a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></strong></em></p></blockquote><blockquote><p><em><strong>Find me on <a href="https://github.com/Abonia1">Github</a></strong></em></p></blockquote><blockquote><p><strong>Visit my technical channel on <a href="https://www.youtube.com/@AboniaSojasingarayar">Youtube</a></strong></p></blockquote><blockquote><p><strong>Support:</strong> <strong><a href="https://www.buymeacoffee.com/abonia">Buy me a Cofee/Chai</a></strong></p></blockquote>]]></content:encoded></item><item><title><![CDATA[LLM Series - Quantization Overview]]></title><description><![CDATA[Enhancing Efficiency While Maintaining Quality]]></description><link>https://aboniasojasingarayar.substack.com/p/llm-series-quantization-overview-1b37c560946b</link><guid isPermaLink="false">https://aboniasojasingarayar.substack.com/p/llm-series-quantization-overview-1b37c560946b</guid><dc:creator><![CDATA[Abonia Sojasingarayar]]></dc:creator><pubDate>Wed, 06 Sep 2023 06:55:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f21e566f-c86d-47e0-a375-4f22e95be1b5_800x600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Enhancing Efficiency While Maintaining Quality</h4><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D-bi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D-bi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png 424w, https://substackcdn.com/image/fetch/$s_!D-bi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png 848w, https://substackcdn.com/image/fetch/$s_!D-bi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png 1272w, https://substackcdn.com/image/fetch/$s_!D-bi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D-bi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D-bi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png 424w, https://substackcdn.com/image/fetch/$s_!D-bi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png 848w, https://substackcdn.com/image/fetch/$s_!D-bi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png 1272w, https://substackcdn.com/image/fetch/$s_!D-bi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12400eb8-07ec-494d-9b27-80db3b0751c4_800x600.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a><figcaption class="image-caption">Quantization&#8202;&#8212;&#8202;Illustration</figcaption></figure></div><p>Quantization, a technique at the forefront of deep learning, is revolutionizing the landscape of neural network deployment. In this article, we delve into the concept of quantization, its types, advantages, and the practical steps to achieve optimal results. As we explore this powerful technique, we&#8217;ll uncover how quantization strikes the delicate balance between computational efficiency and model accuracy.</p><p>Quantization is the process of representing weights, bias and activations in neural networks using lower-precision data types, such as 8-bit integers (int8), instead of the conventional 32-bit floating point (float32) representation. By doing so, it significantly reduces the memory footprint and computational demands during inference, enabling deployment on resource-constrained devices.</p><p>Interested in Parameter Efficient Finetuning. Please do visit&#128071;</p><p><strong><a href="https://medium.com/@abonia/llm-series-parameter-efficient-fine-tuning-e9839fae44ac" title="https://medium.com/@abonia/llm-series-parameter-efficient-fine-tuning-e9839fae44ac">LLM Series&#8202;&#8212;&#8202;Parameter Efficient Fine Tuning</a></strong><a href="https://medium.com/@abonia/llm-series-parameter-efficient-fine-tuning-e9839fae44ac" title="https://medium.com/@abonia/llm-series-parameter-efficient-fine-tuning-e9839fae44ac"><br></a><em><a href="https://medium.com/@abonia/llm-series-parameter-efficient-fine-tuning-e9839fae44ac" title="https://medium.com/@abonia/llm-series-parameter-efficient-fine-tuning-e9839fae44ac">Maximizing Model Performance with Minimal Resources</a></em><a href="https://medium.com/@abonia/llm-series-parameter-efficient-fine-tuning-e9839fae44ac" title="https://medium.com/@abonia/llm-series-parameter-efficient-fine-tuning-e9839fae44ac">medium.com</a></p><h3>Diving into Quantization Types:</h3><h4>Number Representation:</h4><p>For context on how quantization works, explore computer number representation:<br>&#8202;&#8212;&#8202;Integer Representation: Binary representation for unsigned and signed integers.<br>&#8202;&#8212;&#8202;Real Number Representation: Floating-point representation with sign, exponent, and fraction or significand or mantissa.</p><h4>1. Float32 to Float16 Quantization:</h4><p>In this scenario, the transition is from 32-bit floating-point representation to 16-bit floating-point representation. Both data types share the same representation scheme, facilitating a straightforward conversion process. However, compatibility with float16 operations and hardware support is vital for successful implementation.</p><h4>2. Float32 to bfloat16 Quantization:</h4><p>Similar to float16, bfloat16 quantization involves transitioning from 32-bit floating-point to 16-bit floating-point representation, with a specific format known as bfloat16. Bfloat16 offers greater dynamic range compared to float16.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jWAd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jWAd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png 424w, https://substackcdn.com/image/fetch/$s_!jWAd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png 848w, https://substackcdn.com/image/fetch/$s_!jWAd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png 1272w, https://substackcdn.com/image/fetch/$s_!jWAd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jWAd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jWAd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png 424w, https://substackcdn.com/image/fetch/$s_!jWAd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png 848w, https://substackcdn.com/image/fetch/$s_!jWAd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png 1272w, https://substackcdn.com/image/fetch/$s_!jWAd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7e25ef2-9a60-48ac-9af0-ea1f6531b72b_800x337.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Comparison of the float32, bfloat16, and float16 numerical formats&#8202;&#8212;&#8202;<a href="https://www.google.com/url?sa=i&amp;url=https%3A%2F%2Fwww.cerebras.net%2Fmachine-learning%2Fto-bfloat-or-not-to-bfloat-that-is-the-question%2F&amp;psig=AOvVaw1rpP7mUkQTsclplZX5d6pD&amp;ust=1693478291299000&amp;source=images&amp;cd=vfe&amp;opi=89978449&amp;ved=0CAQQjB1qFwoTCLij0bSYhIEDFQAAAAAdAAAAABAf">Cerebras</a></figcaption></figure></div><h4>3. Float32 to Int8 Quantization:</h4><p>This quantization type poses more challenges due to the limited range of representable values in int8 compared to float32. The essence here is to carefully project the float32 value range onto the int8 space to ensure precision is maintained.</p><p>One big challenge for representing weights using lower precision is the smaller numerical range an INT8 can represent, as shown below:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tAaj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tAaj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png 424w, https://substackcdn.com/image/fetch/$s_!tAaj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png 848w, https://substackcdn.com/image/fetch/$s_!tAaj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png 1272w, https://substackcdn.com/image/fetch/$s_!tAaj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tAaj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tAaj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png 424w, https://substackcdn.com/image/fetch/$s_!tAaj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png 848w, https://substackcdn.com/image/fetch/$s_!tAaj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png 1272w, https://substackcdn.com/image/fetch/$s_!tAaj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3622ef40-8e28-4f10-b1bd-f937d2dc3f3c_761x204.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption"><a href="https://randxie.github.io/">Source</a></figcaption></figure></div><h3>Types of Quantization Strategies</h3><blockquote><p><strong>1.Post-Training Quantization</strong></p></blockquote><blockquote><p><strong>2.Quantization-Aware Training</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D_Q8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D_Q8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png 424w, https://substackcdn.com/image/fetch/$s_!D_Q8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png 848w, https://substackcdn.com/image/fetch/$s_!D_Q8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png 1272w, https://substackcdn.com/image/fetch/$s_!D_Q8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D_Q8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D_Q8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png 424w, https://substackcdn.com/image/fetch/$s_!D_Q8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png 848w, https://substackcdn.com/image/fetch/$s_!D_Q8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png 1272w, https://substackcdn.com/image/fetch/$s_!D_Q8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cc97a7c-345d-45ff-afca-318835d8e6c1_800x320.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">QAT&#8202;&#8212;&#8202;PTQ&#8202;&#8212;&#8202;<a href="https://arxiv.org/pdf/2103.13630.pdf">arxiv:2103.13630</a></figcaption></figure></div><h4><strong>1. Post-Training Quantization (PTQ)</strong></h4><p>It involves the quantization of a trained model after the completion of its training phase. By reducing the precision of model parameters, typically from 32-bit floating-point representation to 8-bit integers, PTQ offers alluring benefits such as reduced memory consumption, faster inference times, and improved energy efficiency. However, PTQ often comes at the cost of model accuracy due to the mismatch between the original model and its quantized counterpart.</p><p><strong>GGML vs GPTQ </strong><br>GGML and GPTQ are both quantized models designed to reduce model complexity and computational requirements by using lower-precision model weights. Here&#8217;s a brief comparison of these two approaches:</p><ul><li><p>Optimization Targets: GGML models are optimized for CPU performance, making them faster on CPUs, while GPTQ models are tailored for GPUs, delivering faster inference on GPU hardware.</p></li><li><p>Inference Quality: Inference quality is believed to be similar for both GGML and GPTQ models, but some reports suggest GPTQ might perform slightly lower in specific scenarios.</p></li><li><p>Model Size: GGML models tend to be slightly larger than GPTQ models, which is an important consideration for resource requirements.</p></li><li><p>Compatibility: Both GGML and GPTQ models work seamlessly with Hugging Face Transformers, simplifying their integration into NLP tasks.</p></li></ul><p><strong>Choosing the Right Model:</strong></p><ul><li><p>If you have a CPU without an Nvidia GPU, GGML is recommended.</p></li><li><p>If you have an Nvidia GPU (even if it&#8217;s not the most powerful), GPTQ is a suitable choice.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8QQ7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8QQ7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png 424w, https://substackcdn.com/image/fetch/$s_!8QQ7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png 848w, https://substackcdn.com/image/fetch/$s_!8QQ7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png 1272w, https://substackcdn.com/image/fetch/$s_!8QQ7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8QQ7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8QQ7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png 424w, https://substackcdn.com/image/fetch/$s_!8QQ7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png 848w, https://substackcdn.com/image/fetch/$s_!8QQ7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png 1272w, https://substackcdn.com/image/fetch/$s_!8QQ7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d8363f7-6050-4ad5-aabf-befe0d406331_800x598.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">GGML vs GPTQ&#8202;&#8212;&#8202;Source:1littlecoder</figcaption></figure></div><h4><strong>2. Quantization-Aware Training&nbsp;(QAT)</strong></h4><p>A technique that refines the PTQ model to maintain accuracy even after quantization. Unlike PTQ, where quantization is applied as a separate step, QAT incorporates quantization during the training process itself. By incorporating quantization-related operations (scaling, clipping, and rounding) into the training process, QAT optimizes model weights to mitigate the potential accuracy loss associated with quantization.</p><p>The remarkable aspect of QAT is that it eliminates the need for separate calibration after the training process. The model undergoes calibration as part of training, allowing it to effectively adapt to the quantization constraints. Consequently, the model becomes &#8216;<em><strong>quantization-aware</strong></em>,&#8217; ensuring that accuracy is preserved during real-world inference.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jXLI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jXLI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png 424w, https://substackcdn.com/image/fetch/$s_!jXLI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png 848w, https://substackcdn.com/image/fetch/$s_!jXLI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png 1272w, https://substackcdn.com/image/fetch/$s_!jXLI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jXLI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:null,&quot;width&quot;:null,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jXLI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png 424w, https://substackcdn.com/image/fetch/$s_!jXLI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png 848w, https://substackcdn.com/image/fetch/$s_!jXLI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png 1272w, https://substackcdn.com/image/fetch/$s_!jXLI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda981d9a-bab2-40e6-8184-a44a4b978a15_800x206.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">PTQ vs QAT&#8202;&#8212;&#8202;Q/DQ (quantize then de-quantize)&#8202;&#8212;&#8202;<a href="https://developer.nvidia.com/">Source</a></figcaption></figure></div><p><strong>QAT Mechanism</strong></p><p>QAT employs a novel approach of utilizing &#8216;fake&#8217; quantization modules, marked as Q/DQ (quantize then de-quantize), during training. This enables the model to acclimatize to low-precision weights and account for calculation errors inherent to quantization. The loss function in QAT fine-tunes the model by considering these errors, further enhancing the model&#8217;s ability to maintain accuracy post-quantization.</p><p><em><strong>1. Naive Quantization</strong></em></p><p>Naive quantization involves applying uniform quantization to all operators, leading to a uniform drop in model accuracy. While easy to implement, this method does not account for varying sensitivities of different layers to quantization errors.</p><p><em><strong>2. Hybrid Quantization</strong></em></p><p>Hybrid quantization strikes a balance by quantizing some operators to INT8 precision while leaving others in higher precision (FP16 or FP32). Achieving this balance requires prior knowledge of the model&#8217;s sensitivity to quantization. Despite the challenge of identifying quantization-sensitive layers, hybrid quantization offers better accuracy and latency compared to naive quantization.</p><p><em><strong>3. Selective Quantization</strong></em></p><p>It quantizes specific operators to INT8 precision, employing diverse calibration methods and granularities (per channel or per tensor). This approach accommodates layers that thrive in higher precision due to sensitivity, as well as those that excel with INT8 precision. By offering the flexibility to tailor quantization parameters to different parts of the network, selective quantization maximizes accuracy and minimizes latency simultaneously.</p><h3>&#127919; Advantages of Quantization:</h3><p>Quantization offers several compelling benefits, making it a cornerstone in neural network optimization:</p><p><strong>1. Trimmed Memory Consumption: </strong><br>&nbsp;Quantized models require significantly less memory storage, a crucial advantage for deployment on devices with restricted memory capacity.</p><p><strong>2. Reduced Energy Consumption: </strong><br>&nbsp;Theoretically, quantized models may consume less energy due to reduced data movement and storage operations, contributing to sustainability.</p><p><strong>3. Turbocharged Inference: </strong><br>&nbsp;Integer arithmetic is generally faster than floating-point arithmetic, leading to speedup in operations like matrix multiplications and boosting computational efficiency.</p><p><strong>4. Embedding in Limited Devices:</strong>&nbsp;<br>&nbsp;Many embedded devices only support integer data types. Quantization paves the way for deploying models on such devices that lack native floating-point support.</p><h3>&#10071; Navigating Quantization Challenges:</h3><p>Quantization is not without its challenges:</p><p><strong>1. Overflow and Underflow: </strong><br>&nbsp;Careful scaling and clipping are essential to prevent quantized values from causing overflow or underflow issues.</p><p><strong>2. Symmetric vs. Affine Quantization: </strong><br>&nbsp;The choice between symmetric and affine quantization impacts the arithmetic operations and precision of the quantized model.</p><p><strong>3. Per-Tensor vs. Per-Channel Quantization: </strong><br>&nbsp;The granularity of quantization parameters can be varied to balance accuracy and memory requirements, adding complexity to the process.</p><h3>Practical Implementation Steps:</h3><p>For successful quantization, follow these steps:</p><p><strong>1. Select Quantization-Prone Operators: </strong><br>&nbsp;Identify operators with high computation demands, like matrix multiplications.</p><p><strong>2. Dynamic Quantization Trial: </strong><br>&nbsp;Test dynamic quantization for speed; if satisfactory, stop here.</p><p><strong>3. Static Quantization Experimentation: </strong><br>&nbsp;For improved speed, apply post-training static quantization with observers.</p><p><strong>4. Calibration Technique Choice: </strong><br>&nbsp;Opt for a suitable calibration technique like min-max, moving average min-max, or histogram.</p><p><strong>5. Model Conversion: </strong><br>&nbsp;Remove observers and convert float32 operators to int8 counterparts.</p><p><strong>6. Quantized Model Evaluation: </strong><br>&nbsp;Check if accuracy meets the requirements; if not, consider quantization-aware training.</p><h3><strong>Conclusion:</strong></h3><p>In the realm of large language models (LLMs), the integration of quantization techniques has emerged as a pivotal enabler. Originally devised to enhance efficiency on constrained devices, quantization has now evolved to make fine-tuned LLMs more accessible to diverse users. Quantization emerges as a potent technique in the deep learning area, enabling the deployment of sophisticated models on devices with limited resources. So that its possible for us to efficiently finetune the LLM like Llama with single GPU in classical colab environment. By understanding the nuances of quantization types, challenges, and practical implementation steps, we can harness its power to achieve a harmonious balance between computational efficiency and model accuracy.</p><p>Stay tuned for another interesting article in the LLM series. To receive updates on future articles, please subscribe, keep learning, and keep rocking!</p><h3>&#128736;&#65039; Helpful&nbsp;Resource</h3><ol><li><p><a href="https://arxiv.org/abs/2305.14314">QLoRA: Efficient Finetuning of Quantized LLMs</a></p></li><li><p><a href="https://abvijaykumar.medium.com/fine-tuning-llm-parameter-efficient-fine-tuning-peft-lora-qlora-part-1-571a472612c4">Fine-tuning LLM: Parameter Efficient Fine-tuning (PEFT), LoRA, QLoRA</a></p></li><li><p><a href="https://code4ai.eu/4bit-qlora">4-bit QLoRA</a></p></li><li><p><a href="https://github.com/artidoro/qlora/blob/main/README.md">QLoRA: Efficient Finetuning of Quantized LLMs</a></p></li><li><p><a href="https://pytorch.org/tutorials/intermediate/dynamic_quantization_bert_tutorial.html">PyTorch Quantization</a></p></li></ol><blockquote><p><em><strong>Connect with me on </strong><a href="https://www.linkedin.com/in/aboniasojasingarayar/">Linkedin</a></em></p></blockquote><blockquote><p><em><strong>Find me on </strong><a href="https://github.com/Abonia1">Github</a></em></p></blockquote><blockquote><p><strong>Visit my technical channel on</strong> <a href="http://www.youtube.com/@aboniasojasingarayar3097">Youtube</a></p></blockquote><blockquote><p><em><strong>Support:</strong> <a href="https://www.buymeacoffee.com/abonia">Buy me a Cofee/Chai</a></em></p></blockquote>]]></content:encoded></item></channel></rss>