# LangChain: Framework for LLMs

`Note: I posted this on a different platform too, and there were quite a few people who read it. Hope you like reading this here too.`

## What is LangChain?

It is a framework that helps build LLM-driven applications. One of the problems LangChain tries to solve is RAG or Retrieval Augmented Generation.

## RAG or Retrieval Augmented Generation

RAG is a model architecture utilized in certain LLMs. Once a user sends a query to the LLM, a **retriever** module selects all the related information to the query. The retrieved information is used as context for a **generator** module which **augments** the relevant response.

[![Working of RAG](https://media.dev.to/cdn-cgi/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fg68vds38wc0b3cdagwr0.png align="left")](https://media.dev.to/cdn-cgi/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fg68vds38wc0b3cdagwr0.png)

## Ingestion

Ingestion is about transforming the documents into numbers that computers can understand called **vectors**. The whole process includes:

1. **Loading the source document:** This is done with the help of loaders. The source documents can be PDF, CSV, JSON files or any other web source too.
    
2. **Splitting the document's data into chunks:** This is done with the help of splitters. They split the data into smaller chunks.
    
3. **Converting each of those chunks into vectors:** The chunks are converted to vectors which are known as embeddings, which are then stored to the vector store.
    

[![Process of Ingestion](https://media.dev.to/cdn-cgi/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fh6unihw69yprtrmi4s9a.png align="left")](https://media.dev.to/cdn-cgi/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fh6unihw69yprtrmi4s9a.png)

## What happens when you query the LLM?

Once the data is loaded and stored into the vector store, when a user sends a query, behind the scenes the query gets converted into embeddings. The vector store is checked for similar vectors and is returned. The vectors are then converted into text as a response to the user.

[![bts of querying](https://media.dev.to/cdn-cgi/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzkdo0ps8vr5zj9uri93v.png align="left")](https://media.dev.to/cdn-cgi/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzkdo0ps8vr5zj9uri93v.png)

## The architecture of a chatbot in LangChain

In Langchain, one of the architectures of a chatbot that is followed is shown below. It is structured in such a way that it can handle follow-up questions.

Initially, the chat history and the new question are passed to the LLM. Then, the question is passed to the LLM and asked to create a standalone question. The relevant documents are also retrieved from the vector store. Now, both the relevant documents and standalone question will be used by the LLM to generate a response to the user.

[![Architecture of Chatbot](https://media.dev.to/cdn-cgi/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4xqu1pqt9s13bhxyxccq.png align="left")](https://media.dev.to/cdn-cgi/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4xqu1pqt9s13bhxyxccq.png)

## Technical Terms

A few technical terms associated with LangChain:

* **Embedding:** Vector representations of text, that capture the semantic meaning of the text.
    
* **Runnable:** Allows you to run a set of runnables like model, prompt, parser, etc. in a pre-defined order.
    
* **Streaming:** Allows you to get the output chunk by chunk.
    
* **Batch:** Allows you to process multiple inputs efficiently in a batch or parallelized manner.
    

## Sources

LangChain Documentation: [https://js.langchain.com/docs/get\_started](https://js.langchain.com/docs/get_started)  
LangChain YouTube Channel: [https://youtu.be/AKsfHK\_4tf4?si=ZEIlHWBp4Qfco3O-](https://youtu.be/AKsfHK_4tf4?si=ZEIlHWBp4Qfco3O-)
