diff --git a/docs/html_converter.md b/docs/html_converter.md index 04d45db..3b09fa5 100644 --- a/docs/html_converter.md +++ b/docs/html_converter.md @@ -11,6 +11,7 @@ The `HtmlConverter` utility class provides a mechanism to convert HTML strings i - **Namespace Safety**: Rewrites undeclared prefixed tags (e.g. `o:p` -> `p`) and removes undeclared prefixed attributes (e.g. `foo:bar`) to avoid XML parser `unbound prefix` failures. - **XPath Compatibility**: Strips default XHTML `xmlns` declarations so unprefixed XPath selectors (e.g. `.//div`, `.//h1`) work on converted output. - **Malformed Node Cleanup**: Removes comment and processing-instruction nodes that can serialize into XML-invalid token sequences. +- **Script/Style Removal**: Removes ` + +

Real content after.

+ + +""" + xhtml_output = converter.convert(html_input) + + assert " None: def _remove_problem_nodes(self, doc: Any) -> None: if not hasattr(doc, "xpath"): return - for query in ("//comment()", "//processing-instruction()"): + for query in ( + "//comment()", + "//processing-instruction()", + "//script", + "//style", + ): for node in doc.xpath(query): parent = node.getparent() if parent is not None: