Skip to content

Get all the page content #1037

Description

@jeans11

Hi! First, thanks for this librairie.

Currently, I have a JS app that use jsdom with the plugin readability of Mozilla (that allow to get all the readable content of the page). But for many reasons, I have to rewrite this lib in Java. I'm wondering if HtmtUnit offers a way to get all the readable content of a page.

Cheers

Activity

rbri commented on Oct 6, 2025

@rbri
Member

Hi @jeans11,

this should be possible, in general there are two methods generating the the text content from a node (e.g. the body node).
You can search for the body node and you can call asNormalizedText() or getVisibleText() on the node.

Maybe you will face some problems with this, please report you findings. Will try to support your case as much as possible.
Btw. do you have a public test page?

jeans11 commented on Oct 6, 2025

@jeans11
Author

Hi @rbri

Thank you for the fast answer!

For the moment, I don't have pubic test page yet. I will try the lib with you advice and I will reach you with my result.

rbri commented on Oct 6, 2025

@rbri
Member

After some research, I think i understand what you plan to do. Will at least add some sample content to the home page that might help you a bit with this.

jeans11 commented on Oct 6, 2025

@jeans11
Author

So, after spending few hours to try the lib, I'm really impressed by the capabilities of it!

Very great work @rbri!

The solution to find all the readable text works but a bit more ^^: The lib retrieves all the text of the ads present of the page. I don't know if there is a solution to exclude this kind of content. I can't rely of XPATH because it's no scalable (I can't predict the shape of the page because there are a lots of pages).

rbri commented on Oct 6, 2025

@rbri
Member

@jeans11

have added some words about text extraction to the docu - https://www.htmlunit.org/gettingStarted.html#Extracting_text - feedback welcome

rbri commented on Oct 6, 2025

@rbri
Member

The lib retrieves all the text of the ads present of the page. I don't know if there is a solution to exclude this kind of content. I can't rely of XPATH because it's no scalable (I can't predict the shape of the page because there are a lots of pages).

Now the interesting part starts :-D
Do you have an idea how this was achieved in your existing (js based) solution?

I can think about two possible solutions

  • implement some kind of add blocker (e.g. block all pages not related to the url you like to extract the text from, extract the text) see https://www.htmlunit.org/details.html#Content_blocking for the technical options you have
  • try to manipulate the dom before text extraction. Starting from the page or body element you have more or less all the functions you know from js to walk and change the dom tree. Maybe you only have to iterate over the whole tree and remove the ad-nodes (i guess the tricky part is the destiction between add stuff and real content).

Or you can also subclass the class org.htmlunit.html.serializer.HtmlSerializerNormalizedText and plug in your own implementations. To use your subclass you have to do something like

final AdBlockingHtmlSerializerNormalizedText ser = new AdBlockingHtmlSerializerNormalizedText();
String contentWithoutAds = ser.asText(bodyNode);

Hope that helps.

rbri commented on Oct 6, 2025

@rbri
Member

So, after spending few hours to try the lib, I'm really impressed by the capabilities of it!
Very great work @rbri!

And btw. I guess you use HtmlUnit for some business task - maybe your company can think about sponsoring....

jeans11 commented on Oct 7, 2025

@jeans11
Author

Thanks again @rbri for your response and the code added to the documentation.

The content blocking could improve the parsing of the page. We can retrieve the adblock list and do stuff with it in order to block external ad requests. This is a real possibility but with some drawnbacks compare to Firefox readability (because he strips the ads for us).

So for the moment, I will continue with the JS part.

If one day, we change our mind and make our own solution with HtmlUnit, we sponsoring you with great pleasure.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions