Skip to content

fix: strip script/style elements when converting HTML to XHTML - #40

Open
kisjelica wants to merge 1 commit into
mainfrom
fix/strip-script-style-from-xhtml
Open

kisjelica wants to merge 1 commit into
mainfrom
fix/strip-script-style-from-xhtml

Conversation

@kisjelica

Copy link
Copy Markdown
Contributor

XPath text extraction (e.g. normalize-space() over a content container) returns a node's full string-value, which includes descendant text nodes inside nested <script> and <style> elements. Sites that inject JSON-LD or other markup via client-side JS directly into a content container (as seen on maybelline.com.tr, where BreadcrumbList/HowTo/Article JSON-LD ended up inside schema:articleBody) leak that markup into any graph-sync mapping that extracts text from the container.

HtmlConverter already strips comment and processing-instruction nodes before XPath mappings run; extend the same pass to remove <script> and <style> elements so their text never reaches XPath-based extraction.

XPath text extraction (e.g. normalize-space() over a content container)
returns a node's full string-value, which includes descendant text nodes
inside nested <script> and <style> elements. Sites that inject JSON-LD or
other markup via client-side JS directly into a content container (as
seen on maybelline.com.tr, where BreadcrumbList/HowTo/Article JSON-LD
ended up inside schema:articleBody) leak that markup into any graph-sync
mapping that extracts text from the container.

HtmlConverter already strips comment and processing-instruction nodes
before XPath mappings run; extend the same pass to remove <script> and
<style> elements so their text never reaches XPath-based extraction.
@kisjelica
kisjelica requested a review from rpanfili September 25, 2026 10:01
@rpanfili

Copy link
Copy Markdown
Contributor

Hey Rastko, the HtmlConverter responsibility

Convert an HTML string to a valid XHTML string.

seems generic enough that this change shouldn’t necessarily be applied there. It seems to me that it should be the responsibility of the consumer of that XHTML to eventually strip/skip undesired tags.

What’s your reading on this?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants