Building the First AI App in .NET: Step 3 – From Keywords to Semantic Search

Intro In the previous step we made the application a little more useful by introducing a very simple form of retrieval. Instead of sending every document to the LLM, we split the documents into chunks and tried to find the ones that looked relevant to the question. But there is an obvious problem with that approach which is we were essentially just counting words. That works when the question and the document use the same terminology, but it breaks very easily when the same idea is expressed differently. Imagine that one of our documents says:Authentication uses JWT tokens. And the user asks:How do users log into the system? This is where embeddings become useful. There isn’t much overlap between the above two sentences if we look at the actual words. A simple keyword search could easily miss the relevant document even though a human would immediately understand that they are talking about the same thing. This is where embeddings become useful. As usual, here is the code for what we are about to do: https://github.com/genoiucosmin/AIApp1/tree/Step3 Moving from words to meaning An embedding is a numerical representation of text. Instead of treating a sentence simply as a collection of words, an…

Building the First AI App in .NET – Step 2: From Grounding to Retrieval

Introduction In Step 1, we built a simple grounded LLM application. The application could answer questions using our documents, but there was a fundamental problem: we sent all documents to the model for every question. That works for a handful of documents. It doesn’t work for a real knowledge base. So the next question is: How do we give the LLM only the information it actually needs? This leads us to retrieval and the moment we learn about RAG. As usual, you can find the code for this example here: https://github.com/genoiucosmin/AIApp1/tree/Step2 From all documents to relevant documents Step 1 Step 2 Let’s explain why this is cheaper, faster and more scalable conceptually. Chunking and our first retrieval mechanism In the previous example we sent in all documents to the LLM. This is clearly not something scalable, so we need to define a way of sending only the relevant data to the LLM. This is actually RAG – and we will implement a very primitive version of it, by just splitting the documents we had into chunks of 300 characters. Then we find out how many times each word from our question appears in each chunk. We then select the top…