Community[minor]: Update VDMS vectorstore (#23729)

**Description:** - This PR exposes some functions in VDMS vectorstore, updates VDMS related notebooks, updates tests, and upgrade version of VDMS (>=0.0.20) **Issue:** N/A **Dependencies:** - Update vdms>=0.0.20
2025-09-01 11:02:37 +00:00 · 2024-07-25 19:13:04 -07:00
parent 703491e824
commit 69eacaa887
8 changed files with 627 additions and 350 deletions
--- a/cookbook/README.md
+++ b/cookbook/README.md
@@ -36,6 +36,7 @@ Notebook | Description
 [llm_symbolic_math.ipynb](https://github.com/langchain-ai/langchain/tree/master/cookbook/llm_symbolic_math.ipynb) | Solve algebraic equations with the help of llms (language learning models) and sympy, a python library for symbolic mathematics.
 [meta_prompt.ipynb](https://github.com/langchain-ai/langchain/tree/master/cookbook/meta_prompt.ipynb) | Implement the meta-prompt concept, which is a method for building self-improving agents that reflect on their own performance and modify their instructions accordingly.
 [multi_modal_output_agent.ipynb](https://github.com/langchain-ai/langchain/tree/master/cookbook/multi_modal_output_agent.ipynb) | Generate multi-modal outputs, specifically images and text.
+[multi_modal_RAG_vdms.ipynb](https://github.com/langchain-ai/langchain/tree/master/cookbook/multi_modal_RAG_vdms.ipynb) | Perform retrieval-augmented generation (rag) on documents including text and images, using unstructured for parsing, Intel's Visual Data Management System (VDMS) as the vectorstore, and chains.
 [multi_player_dnd.ipynb](https://github.com/langchain-ai/langchain/tree/master/cookbook/multi_player_dnd.ipynb) | Simulate multi-player dungeons & dragons games, with a custom function determining the speaking schedule of the agents.
 [multiagent_authoritarian.ipynb](https://github.com/langchain-ai/langchain/tree/master/cookbook/multiagent_authoritarian.ipynb) | Implement a multi-agent simulation where a privileged agent controls the conversation, including deciding who speaks and when the conversation ends, in the context of a simulated news network.
 [multiagent_bidding.ipynb](https://github.com/langchain-ai/langchain/tree/master/cookbook/multiagent_bidding.ipynb) | Implement a multi-agent simulation where agents bid to speak, with the highest bidder speaking next, demonstrated through a fictitious presidential debate example.
--- a/cookbook/multi_modal_RAG_vdms.ipynb
+++ b/cookbook/multi_modal_RAG_vdms.ipynb
@@ -18,26 +18,7 @@
    "* Use of multimodal embeddings (such as [CLIP](https://openai.com/research/clip)) to embed images and text\n",
    "* Use of [VDMS](https://github.com/IntelLabs/vdms/blob/master/README.md) as a vector store with support for multi-modal\n",
    "* Retrieval of both images and text using similarity search\n",
-    "* Passing raw images and text chunks to a multimodal LLM for answer synthesis \n",
-    "\n",
-    "\n",
-    "## Packages\n",
-    "\n",
-    "For `unstructured`, you will also need `poppler` ([installation instructions](https://pdf2image.readthedocs.io/en/latest/installation.html)) and `tesseract` ([installation instructions](https://tesseract-ocr.github.io/tessdoc/Installation.html)) in your system."
-   ]
-  },
-  {
-   "cell_type": "code",
-   "execution_count": 1,
-   "id": "febbc459-ebba-4c1a-a52b-fed7731593f8",
-   "metadata": {},
-   "outputs": [],
-   "source": [
-    "# (newest versions required for multi-modal)\n",
-    "! pip install --quiet -U vdms langchain-experimental\n",
-    "\n",
-    "# lock to 0.10.19 due to a persistent bug in more recent versions\n",
-    "! pip install --quiet pdf2image \"unstructured[all-docs]==0.10.19\" pillow pydantic lxml open_clip_torch"
+    "* Passing raw images and text chunks to a multimodal LLM for answer synthesis "
   ]
  },
  {
@@ -53,7 +34,7 @@
  },
  {
   "cell_type": "code",
-   "execution_count": 3,
+   "execution_count": 1,
   "id": "5f483872",
   "metadata": {},
   "outputs": [
@@ -61,8 +42,7 @@
     "name": "stdout",
     "output_type": "stream",
     "text": [
-      "docker: Error response from daemon: Conflict. The container name \"/vdms_rag_nb\" is already in use by container \"0c19ed281463ac10d7efe07eb815643e3e534ddf24844357039453ad2b0c27e8\". You have to remove (or rename) that container to be able to reuse that name.\n",
-      "See 'docker run --help'.\n"
+      "a1b9206b08ef626e15b356bf9e031171f7c7eb8f956a2733f196f0109246fe2b\n"
     ]
    }
   ],
@@ -75,9 +55,32 @@
    "vdms_client = VDMS_Client(port=55559)"
   ]
  },
+  {
+   "cell_type": "markdown",
+   "id": "2498a0a1",
+   "metadata": {},
+   "source": [
+    "## Packages\n",
+    "\n",
+    "For `unstructured`, you will also need `poppler` ([installation instructions](https://pdf2image.readthedocs.io/en/latest/installation.html)) and `tesseract` ([installation instructions](https://tesseract-ocr.github.io/tessdoc/Installation.html)) in your system."
+   ]
+  },
  {
   "cell_type": "code",
-   "execution_count": null,
+   "execution_count": 2,
+   "id": "febbc459-ebba-4c1a-a52b-fed7731593f8",
+   "metadata": {},
+   "outputs": [],
+   "source": [
+    "! pip install --quiet -U vdms langchain-experimental\n",
+    "\n",
+    "# lock to 0.10.19 due to a persistent bug in more recent versions\n",
+    "! pip install --quiet pdf2image \"unstructured[all-docs]==0.10.19\" pillow pydantic lxml open_clip_torch"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": 3,
   "id": "78ac6543",
   "metadata": {},
   "outputs": [],
@@ -95,14 +98,9 @@
    "\n",
    "### Partition PDF text and images\n",
    "  \n",
-    "Let's look at an example pdf containing interesting images.\n",
+    "Let's use famous photographs from the PDF version of Library of Congress Magazine in this example.\n",
    "\n",
-    "Famous photographs from library of congress:\n",
-    "\n",
-    "* https://www.loc.gov/lcm/pdf/LCM_2020_1112.pdf\n",
-    "* We'll use this as an example below\n",
-    "\n",
-    "We can use `partition_pdf` below from [Unstructured](https://unstructured-io.github.io/unstructured/introduction.html#key-concepts) to extract text and images."
+    "We can use `partition_pdf` from [Unstructured](https://unstructured-io.github.io/unstructured/introduction.html#key-concepts) to extract text and images."
   ]
  },
  {
@@ -116,8 +114,8 @@
    "\n",
    "import requests\n",
    "\n",
-    "# Folder with pdf and extracted images\n",
-    "datapath = Path(\"./multimodal_files\").resolve()\n",
+    "# Folder to store pdf and extracted images\n",
+    "datapath = Path(\"./data/multimodal_files\").resolve()\n",
    "datapath.mkdir(parents=True, exist_ok=True)\n",
    "\n",
    "pdf_url = \"https://www.loc.gov/lcm/pdf/LCM_2020_1112.pdf\"\n",
@@ -174,14 +172,8 @@
   "source": [
    "## Multi-modal embeddings with our document\n",
    "\n",
-    "We will use [OpenClip multimodal embeddings](https://python.langchain.com/docs/integrations/text_embedding/open_clip).\n",
-    "\n",
-    "We use a larger model for better performance (set in `langchain_experimental.open_clip.py`).\n",
-    "\n",
-    "```\n",
-    "model_name = \"ViT-g-14\"\n",
-    "checkpoint = \"laion2b_s34b_b88k\"\n",
-    "```"
+    "In this section, we initialize the VDMS vector store for both text and images. For better performance, we use model `ViT-g-14` from [OpenClip multimodal embeddings](https://python.langchain.com/docs/integrations/text_embedding/open_clip).\n",
+    "The images are stored as base64 encoded strings with `vectorstore.add_images`.\n"
   ]
  },
  {
@@ -200,9 +192,7 @@
    "vectorstore = VDMS(\n",
    "    client=vdms_client,\n",
    "    collection_name=\"mm_rag_clip_photos\",\n",
-    "    embedding_function=OpenCLIPEmbeddings(\n",
-    "        model_name=\"ViT-g-14\", checkpoint=\"laion2b_s34b_b88k\"\n",
-    "    ),\n",
+    "    embedding=OpenCLIPEmbeddings(model_name=\"ViT-g-14\", checkpoint=\"laion2b_s34b_b88k\"),\n",
    ")\n",
    "\n",
    "# Get image URIs with .jpg extension only\n",
@@ -233,7 +223,7 @@
   "source": [
    "## RAG\n",
    "\n",
-    "`vectorstore.add_images` will store / retrieve images as base64 encoded strings."
+    "Here we define helper functions for image results."
   ]
  },
  {
@@ -392,7 +382,8 @@
   "id": "1566096d-97c2-4ddc-ba4a-6ef88c525e4e",
   "metadata": {},
   "source": [
-    "## Test retrieval and run RAG"
+    "## Test retrieval and run RAG\n",
+    "Now let's query for a `woman with children` and retrieve the top results."
   ]
  },
  {
@@ -452,6 +443,14 @@
    "        print(doc.page_content)"
   ]
  },
+  {
+   "cell_type": "markdown",
+   "id": "15e9b54d",
+   "metadata": {},
+   "source": [
+    "Now let's use the `multi_modal_rag_chain` to process the same query and display the response."
+   ]
+  },
  {
   "cell_type": "code",
   "execution_count": 11,
@@ -462,10 +461,10 @@
     "name": "stdout",
     "output_type": "stream",
     "text": [
-      "1. Detailed description of the visual elements in the image: The image features a woman with children, likely a mother and her family, standing together outside. They appear to be poor or struggling financially, as indicated by their attire and surroundings.\n",
-      "2. Historical and cultural context of the image: The photo was taken in 1936 during the Great Depression, when many families struggled to make ends meet. Dorothea Lange, a renowned American photographer, took this iconic photograph that became an emblem of poverty and hardship experienced by many Americans at that time.\n",
-      "3. Interpretation of the image's symbolism and meaning: The image conveys a sense of unity and resilience despite adversity. The woman and her children are standing together, displaying their strength as a family unit in the face of economic challenges. The photograph also serves as a reminder of the importance of empathy and support for those who are struggling.\n",
-      "4. Connections between the image and the related text: The text provided offers additional context about the woman in the photo, her background, and her feelings towards the photograph. It highlights the historical backdrop of the Great Depression and emphasizes the significance of this particular image as a representation of that time period.\n"
+      " The image depicts a woman with several children. The woman appears to be of Cherokee heritage, as suggested by the text provided. The image is described as having been initially regretted by the subject, Florence Owens Thompson, due to her feeling that it did not accurately represent her leadership qualities.\n",
+      "The historical and cultural context of the image is tied to the Great Depression and the Dust Bowl, both of which affected the Cherokee people in Oklahoma. The photograph was taken during this period, and its subject, Florence Owens Thompson, was a leader within her community who worked tirelessly to help those affected by these crises.\n",
+      "The image's symbolism and meaning can be interpreted as a representation of resilience and strength in the face of adversity. The woman is depicted with multiple children, which could signify her role as a caregiver and protector during difficult times.\n",
+      "Connections between the image and the related text include Florence Owens Thompson's leadership qualities and her regretted feelings about the photograph. Additionally, the mention of Dorothea Lange, the photographer who took this photo, ties the image to its historical context and the broader narrative of the Great Depression and Dust Bowl in Oklahoma. \n"
     ]
    }
   ],
@@ -492,14 +491,6 @@
   "source": [
    "! docker kill vdms_rag_nb"
   ]
-  },
-  {
-   "cell_type": "code",
-   "execution_count": null,
-   "id": "8ba652da",
-   "metadata": {},
-   "outputs": [],
-   "source": []
  }
 ],
 "metadata": {
@@ -518,7 +509,7 @@
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
-   "version": "3.10.13"
+   "version": "3.11.9"
  }
 },
 "nbformat": 4,