Repository navigation
Get all the page content #1037
Description
Activity
Hi @jeans11,
this should be possible, in general there are two methods generating the the text content from a node (e.g. the body node).
You can search for the body node and you can call asNormalizedText() or getVisibleText() on the node.
Maybe you will face some problems with this, please report you findings. Will try to support your case as much as possible.
Btw. do you have a public test page?
Hi @rbri
Thank you for the fast answer!
For the moment, I don't have pubic test page yet. I will try the lib with you advice and I will reach you with my result.
After some research, I think i understand what you plan to do. Will at least add some sample content to the home page that might help you a bit with this.
So, after spending few hours to try the lib, I'm really impressed by the capabilities of it!
Very great work @rbri!
The solution to find all the readable text works but a bit more ^^: The lib retrieves all the text of the ads present of the page. I don't know if there is a solution to exclude this kind of content. I can't rely of XPATH because it's no scalable (I can't predict the shape of the page because there are a lots of pages).
have added some words about text extraction to the docu - https://www.htmlunit.org/gettingStarted.html#Extracting_text - feedback welcome
The lib retrieves all the text of the ads present of the page. I don't know if there is a solution to exclude this kind of content. I can't rely of XPATH because it's no scalable (I can't predict the shape of the page because there are a lots of pages).
Now the interesting part starts :-D
Do you have an idea how this was achieved in your existing (js based) solution?
I can think about two possible solutions
- implement some kind of add blocker (e.g. block all pages not related to the url you like to extract the text from, extract the text) see https://www.htmlunit.org/details.html#Content_blocking for the technical options you have
- try to manipulate the dom before text extraction. Starting from the page or body element you have more or less all the functions you know from js to walk and change the dom tree. Maybe you only have to iterate over the whole tree and remove the ad-nodes (i guess the tricky part is the destiction between add stuff and real content).
Or you can also subclass the class org.htmlunit.html.serializer.HtmlSerializerNormalizedText and plug in your own implementations. To use your subclass you have to do something like
final AdBlockingHtmlSerializerNormalizedText ser = new AdBlockingHtmlSerializerNormalizedText();
String contentWithoutAds = ser.asText(bodyNode);
Hope that helps.
So, after spending few hours to try the lib, I'm really impressed by the capabilities of it!
Very great work @rbri!
And btw. I guess you use HtmlUnit for some business task - maybe your company can think about sponsoring....
Thanks again @rbri for your response and the code added to the documentation.
The content blocking could improve the parsing of the page. We can retrieve the adblock list and do stuff with it in order to block external ad requests. This is a real possibility but with some drawnbacks compare to Firefox readability (because he strips the ads for us).
So for the moment, I will continue with the JS part.
If one day, we change our mind and make our own solution with HtmlUnit, we sponsoring you with great pleasure.
Hi! First, thanks for this librairie.
Currently, I have a JS app that use
jsdomwith the pluginreadabilityof Mozilla (that allow to get all the readable content of the page). But for many reasons, I have to rewrite this lib in Java. I'm wondering if HtmtUnit offers a way to get all the readable content of a page.Cheers