fix import error of bs4 (#1952)

Ran into a broken build if bs4 wasn't installed in the project.

Minor tweak to follow the other doc loaders optional package-loading
conventions.

Also updated html docs to include reference to this new html loader.

side note: Should there be 2 different html-to-text document loaders?
This new one only handles local files, while the existing unstructured
html loader handles HTML from local and remote. So it seems like the
improvement was adding the title to the metadata, which is useful but
could also be added to `html.py`
This commit is contained in:
Tim Asp
2023-03-23 21:56:13 -07:00
committed by GitHub
parent 8990122d5d
commit 030ce9f506
3 changed files with 59 additions and 7 deletions

View File

@@ -1,5 +1,8 @@
<!DOCTYPE html>
<html>
<head>
<title>Test Title</title>
</head>
<body>
<h1>My First Heading</h1>