Repository navigation
encoding problem #358
Description
Activity
Can you please provide a bit more details - or better some minimal sample code that shows the porblem.
Hello, I'll try to describe the problem as it happens occasionally. For example, the encoding of a web page is set to: < mate content = "text / htmlcharset = GB2312" HTTP equiv = "content type" >, and then I add the following in line 123 of HtmlUnitNekoHtmlParser class in the crawler project:
if (encoding.equals("GB2312")){encoding = "GBK";}
When I use webclient.getPage(URL) to get the HtmlPage and parse it. Sometimes the Chinese on the page is garbled, sometimes not.There are about tens of thousands of requests.
Or I want to use GBK code to parse the page, what should I do?
Two points,
- GB2312 - GBK
- sometimes the result is garbage
Regarding 1.
Because i'm from Germany i only have limited knowledge about all the Chinese codepage stuff. But i did some research and it looks like the jdk handles GB2312 different than browsers. Because of this i like to write a unit test to check the behavior of real browsers and add a fix to HtmlUnit (if needed). But i think i need your help for writing this test - do you like to help here?
Regarding 2.
Sounds like a multi threading issue. Question: do you get this always for the same page?
Thank you very much for your help. I changed the new version yesterday, and then optimized the program, there is no garbled code at present. If there is a follow-up I can catch the error, and then reply.
Many thanks for pointing at this - was able to write some test cases and have added a fix.
As a side effect a bunch of other existing tests are working now also.
You can try the latest snapshot - and a release is hopefully available soon.
Again thanks and have fun using HtmlUnit.
If the page code is GBK, after calling getpage (URL), occasionally Chinese garbled will appear. How do I set it?