Skip to content

Return NA port from url_parse() for URLs that can't be parsed - #481

Open
taekop wants to merge 2 commits into
r-lib:mainfrom
taekop:fix-url-parse-invalid
Open

taekop wants to merge 2 commits into
r-lib:mainfrom
taekop:fix-url-parse-invalid

Conversation

@taekop

@taekop taekop commented Oct 5, 2026 •

Copy link
Copy Markdown

url_parse() skips the row when xmlParseURI() returns NULL, but the port vector is allocated without being initialised, so a URL libxml2 can't parse (e.g. one with a space or a non-ASCII character) came back with a garbage port such as -541335376. Such rows now get port = NA, matching the empty strings in the other columns.

Closes #442

@jeroen

jeroen commented Oct 6, 2026

Copy link
Copy Markdown
Member

Does anyone use this URL parser? If it is really an issue, it would be better to zero-out the port vector, rather than guessing to fix the broken URL. Your current solution also escapes characters that should not be escaped so it is not a good solution.

@taekop taekop changed the title Parse URLs with unescaped characters in url_parse() Return NA port from url_parse() for URLs that can't be parsed Oct 6, 2026
@taekop

taekop commented Oct 6, 2026

Copy link
Copy Markdown
Author

Agreed, I dropped the retry. The PR now only sets port to NA for URLs that fail to parse, so the other columns stay empty as before. The only user report I know of is #442.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

url_parse doesn't work with URL containing non-ASCII characters

2 participants