Update dependency org.jsoup:jsoup to v1.23.1 #73

Open
renovate wants to merge 1 commits from renovate/org.jsoup-jsoup-1.x into main
Collaborator

This PR contains the following updates:

Package Type Update Change
org.jsoup:jsoup (source) compile minor 1.22.11.23.1

Release Notes

jhy/jsoup (org.jsoup:jsoup)

v1.23.1

Improvements
  • Reduced retained memory when parsing with source position tracking enabled (Parser#setTrackPosition(true)). Source ranges are now stored in compact parser-owned span records instead of node and attribute user data, and Position objects are created lazily when source ranges are read. This cuts tracked DOM retained size by about 50-60% on representative benchmark documents, while keeping Node#sourceRange(), Element#endSourceRange(), and Attribute#sourceRange() behavior intact. #​2498
  • Added Element#classList(), an immutable snapshot of an element's class names in attribute order. Use hasClass() when you just need to test for one class, classList() when you want to read or iterate classes without needing a mutable result, and classNames() when you want the existing mutable, deduplicated set that can be written back with classNames(Set). The class APIs now share an HTML-whitespace scanner, which also makes classNames() faster and lighter on allocation, especially when walking many elements without class names. #​2500
  • Aligned HTML parser scope classification with the current HTML spec for select, foreignObject, and template. #​2501
  • Simplified the HTML tree builder's scope, implied-end-tag, and special-element checks by caching parser-only options on Tag. That improves HTML parser throughput by about 10% on small inputs and up to about 30% on larger inputs in the benchmark fixtures. #​2502
  • Improved HTML parser throughput stability by making hot tokeniser scan paths compile more predictably. #​2507
  • <noscript> fallback markup is now parsed into an inspectable DOM subtree in both the document head and body. The fallback acts as a contained parsing island, so malformed markup cannot disrupt the surrounding document structure, while normal HTML tokenization still applies within it. This also improves round-trip serialization. #​2537
  • Improved redirect credential handling as a defense-in-depth measure: explicit authorization headers and request cookies are no longer forwarded across origins, reducing exposure through open redirects and aligning with HTTP guidance. Cookies managed by a CookieStore continue to follow their configured scope. #​2540
  • Elements can now append their outer HTML, including their own tags, directly to an Appendable with Node#outerHtml(Appendable), without first creating a String. This complements Element#html(Appendable), which appends inner HTML only. #​2532
  • Aligned CDATA tokenization with the HTML spec: CDATA syntax in HTML content is parsed as a bogus comment, while it remains supported in SVG, MathML, and XML. Also improved namespace-aware fragment parsing so SVG and MathML contexts, HTML integration points, and context-sensitive tokenizer states are handled correctly. #​2542
  • When using the optional re2j regular expression engine, stack overflows caused by complex selector patterns are now normalized to a ValidationException with a Pattern complexity error message. #​2548
Bug Fixes
  • Fixed HTML parsing of mixed-case RCDATA end tags after tag-shaped text. For example, <title><p>Foo</TiTLE> and <textarea><img src=x></TeXtArEa> now keep the tag-shaped content as text instead of promoting it to markup. #​2503
  • Fixed W3CDom XML conversion so plain XML elements don't serialize with the reserved XML namespace as the default namespace. Explicit XML namespaces and xml:* attributes are still preserved. #​2504
  • Preserve control characters in parsed tag names #​2538
  • Updated HTTP redirects to follow the specification: 307 and 308 preserve the request method and content, 301 and 302 only change POST to GET, and Location is followed only for 301, 302, 303, 307, and 308 responses. Streamed request bodies are not buffered; if an automatic redirect requires replaying one, execution fails, so the caller can resend with a fresh stream. #​2540
  • Corrected the Cleaner's same-site link detection to compare hostnames rather than URL prefixes when applying rel=nofollow. #​2543
Build Changes
  • Cleaned up the Maven build for the multi-release JAR so Java 8 and Java 11+ sources compile as separate source sets. This avoids spurious Java 8 compiler warnings from newer-language overlay sources, keeps long-running parser checks behind an explicit profile, and preserves the same published artifacts and runtime behavior.
  • Improved parallelism and tuned timing in our integration tests, so that a full mvn clean verify drops from ~ 1m18s to ~ 21 seconds.

v1.22.2

Improvements
  • Expanded and clarified NodeTraversor support for in-place DOM rewrites during NodeVisitor.head(). Current-node edits such as remove, replace, and unwrap now recover more predictably, while traversal stays within the original root subtree. This makes single-pass tree cleanup and normalization visitors easier to write, for example when unwrapping presentational elements or replacing text nodes as you walk the DOM. #​2472
  • Documentation: clarified that a configured Cleaner may be reused across concurrent threads, and that shared Safelist instances should not be mutated while in use. #​2473
  • Updated the default HTML TagSet for current HTML elements: added dialog, search, picture, and slot; made ins, del, button, audio, video, and canvas inline by default (Tag#isInline(), aligned to phrasing content in the spec); and added readable Element.text() boundaries for controls and embedded objects via the new Tag.TextBoundary option. This improves pretty-printing and keeps normalized text from running adjacent words together. #​2493
Bug Fixes
  • Android (R8/ProGuard): added a rule to ignore the optional re2j dependency when not present. #​2459
  • Fixed a NodeTraversor regression in 1.21.2 where removing or replacing the current node during head() could revisit the replacement node and loop indefinitely. The traversal docs now also clarify which inserted nodes are visited in the current pass. #​2472
  • Parsing during charset sniffing no longer fails if an advisory available() call throws IOException, as seen on JDK 8 HttpURLConnection. #​2474
  • Cleaner no longer makes relative URL attributes in the input document absolute when cleaning or validating a Document. URL normalization now applies only to the cleaned output, and Safelist.isSafeAttribute() is side effect free. #​2475
  • Cleaner no longer duplicates enforced attributes when the input Document preserves attribute case. A case-variant source attribute is now replaced by the enforced attribute in the cleaned output. #​2476
  • If a per-request SOCKS proxy is configured, jsoup now avoids using the JDK HttpClient, because the JDK would silently ignore that proxy and attempt to connect directly. Those requests now fall back to the legacy HttpURLConnection transport instead, which does support SOCKS. #​2468
  • Connection.Response.streamParser() and DataUtil.streamParser(Path, ...) could fail on small inputs without a declared charset, if the initial 5 KB charset sniff fully consumed the input and closed it before the stream parse began. #​2483
  • In XML mode, doctypes with an internal subset, such as <!DOCTYPE root [<!ENTITY name "value">]>, now round-trip correctly. The subset is preserved as raw text only; entities are not expanded and external DTDs are not loaded. #​2486
Build Changes
  • Migrated the integration test server from Jetty to Netty, which actively maintains support for our minimum JDK target (8). #​2491

Configuration

📅 Schedule: (UTC)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about this update again.


  • If you want to rebase/retry this PR, check this box

This PR has been generated by Mend Renovate CLI.

This PR contains the following updates: | Package | Type | Update | Change | |---|---|---|---| | [org.jsoup:jsoup](https://jsoup.org/) ([source](https://github.com/jhy/jsoup)) | compile | minor | `1.22.1` → `1.23.1` | --- ### Release Notes <details> <summary>jhy/jsoup (org.jsoup:jsoup)</summary> ### [`v1.23.1`](https://github.com/jhy/jsoup/blob/HEAD/CHANGES.md#1231-2026-Jul-30) ##### Improvements - Reduced retained memory when parsing with source position tracking enabled (`Parser#setTrackPosition(true)`). Source ranges are now stored in compact parser-owned span records instead of node and attribute user data, and `Position` objects are created lazily when source ranges are read. This cuts tracked DOM retained size by about 50-60% on representative benchmark documents, while keeping `Node#sourceRange()`, `Element#endSourceRange()`, and `Attribute#sourceRange()` behavior intact. [#&#8203;2498](https://github.com/jhy/jsoup/pull/2498) - Added `Element#classList()`, an immutable snapshot of an element's class names in attribute order. Use `hasClass()` when you just need to test for one class, `classList()` when you want to read or iterate classes without needing a mutable result, and `classNames()` when you want the existing mutable, deduplicated set that can be written back with `classNames(Set)`. The class APIs now share an HTML-whitespace scanner, which also makes `classNames()` faster and lighter on allocation, especially when walking many elements without class names. [#&#8203;2500](https://github.com/jhy/jsoup/pull/2500) - Aligned HTML parser scope classification with the current HTML spec for `select`, `foreignObject`, and `template`. [#&#8203;2501](https://github.com/jhy/jsoup/issues/2501) - Simplified the HTML tree builder's scope, implied-end-tag, and special-element checks by caching parser-only options on Tag. That improves HTML parser throughput by about 10% on small inputs and up to about 30% on larger inputs in the benchmark fixtures. [#&#8203;2502](https://github.com/jhy/jsoup/issues/2502) - Improved HTML parser throughput stability by making hot tokeniser scan paths compile more predictably. [#&#8203;2507](https://github.com/jhy/jsoup/pull/2507) - `<noscript>` fallback markup is now parsed into an inspectable DOM subtree in both the document head and body. The fallback acts as a contained parsing island, so malformed markup cannot disrupt the surrounding document structure, while normal HTML tokenization still applies within it. This also improves round-trip serialization. [#&#8203;2537](https://github.com/jhy/jsoup/pull/2537) - Improved redirect credential handling as a defense-in-depth measure: explicit authorization headers and request cookies are no longer forwarded across origins, reducing exposure through open redirects and aligning with HTTP guidance. Cookies managed by a `CookieStore` continue to follow their configured scope. [#&#8203;2540](https://github.com/jhy/jsoup/pull/2540) - Elements can now append their outer HTML, including their own tags, directly to an `Appendable` with `Node#outerHtml(Appendable)`, without first creating a `String`. This complements `Element#html(Appendable)`, which appends inner HTML only. [#&#8203;2532](https://github.com/jhy/jsoup/issues/2532) - Aligned CDATA tokenization with the HTML spec: CDATA syntax in HTML content is parsed as a bogus comment, while it remains supported in SVG, MathML, and XML. Also improved namespace-aware fragment parsing so SVG and MathML contexts, HTML integration points, and context-sensitive tokenizer states are handled correctly. [#&#8203;2542](https://github.com/jhy/jsoup/issues/2542) - When using the optional `re2j` regular expression engine, stack overflows caused by complex selector patterns are now normalized to a `ValidationException` with a `Pattern complexity error` message. [#&#8203;2548](https://github.com/jhy/jsoup/issues/2548) ##### Bug Fixes - Fixed HTML parsing of mixed-case RCDATA end tags after tag-shaped text. For example, `<title><p>Foo</TiTLE>` and `<textarea><img src=x></TeXtArEa>` now keep the tag-shaped content as text instead of promoting it to markup. [#&#8203;2503](https://github.com/jhy/jsoup/issues/2503) - Fixed `W3CDom` XML conversion so plain XML elements don't serialize with the reserved XML namespace as the default namespace. Explicit XML namespaces and `xml:*` attributes are still preserved. [#&#8203;2504](https://github.com/jhy/jsoup/issues/2504) - Preserve control characters in parsed tag names [#&#8203;2538](https://github.com/jhy/jsoup/issues/2538) - Updated HTTP redirects to follow the specification: 307 and 308 preserve the request method and content, 301 and 302 only change POST to GET, and `Location` is followed only for 301, 302, 303, 307, and 308 responses. Streamed request bodies are not buffered; if an automatic redirect requires replaying one, execution fails, so the caller can resend with a fresh stream. [#&#8203;2540](https://github.com/jhy/jsoup/pull/2540) - Corrected the Cleaner's same-site link detection to compare hostnames rather than URL prefixes when applying `rel=nofollow`. [#&#8203;2543](https://github.com/jhy/jsoup/issues/2543) ##### Build Changes - Cleaned up the Maven build for the multi-release JAR so Java 8 and Java 11+ sources compile as separate source sets. This avoids spurious Java 8 compiler warnings from newer-language overlay sources, keeps long-running parser checks behind an explicit profile, and preserves the same published artifacts and runtime behavior. - Improved parallelism and tuned timing in our integration tests, so that a full `mvn clean verify` drops from \~ 1m18s to \~ 21 seconds. ### [`v1.22.2`](https://github.com/jhy/jsoup/blob/HEAD/CHANGES.md#1222-2026-Apr-20) ##### Improvements - Expanded and clarified `NodeTraversor` support for in-place DOM rewrites during `NodeVisitor.head()`. Current-node edits such as `remove`, `replace`, and `unwrap` now recover more predictably, while traversal stays within the original root subtree. This makes single-pass tree cleanup and normalization visitors easier to write, for example when unwrapping presentational elements or replacing text nodes as you walk the DOM. [#&#8203;2472](https://github.com/jhy/jsoup/issues/2472) - Documentation: clarified that a configured `Cleaner` may be reused across concurrent threads, and that shared `Safelist` instances should not be mutated while in use. [#&#8203;2473](https://github.com/jhy/jsoup/issues/2473) - Updated the default HTML `TagSet` for current HTML elements: added `dialog`, `search`, `picture`, and `slot`; made `ins`, `del`, `button`, `audio`, `video`, and `canvas` inline by default (`Tag#isInline()`, aligned to phrasing content in the spec); and added readable `Element.text()` boundaries for controls and embedded objects via the new `Tag.TextBoundary` option. This improves pretty-printing and keeps normalized text from running adjacent words together. [#&#8203;2493](https://github.com/jhy/jsoup/pull/2493) ##### Bug Fixes - Android (R8/ProGuard): added a rule to ignore the optional `re2j` dependency when not present. [#&#8203;2459](https://github.com/jhy/jsoup/issues/2459) - Fixed a `NodeTraversor` regression in 1.21.2 where removing or replacing the current node during `head()` could revisit the replacement node and loop indefinitely. The traversal docs now also clarify which inserted nodes are visited in the current pass. [#&#8203;2472](https://github.com/jhy/jsoup/issues/2472) - Parsing during charset sniffing no longer fails if an advisory `available()` call throws `IOException`, as seen on JDK 8 `HttpURLConnection`. [#&#8203;2474](https://github.com/jhy/jsoup/issues/2474) - `Cleaner` no longer makes relative URL attributes in the input document absolute when cleaning or validating a `Document`. URL normalization now applies only to the cleaned output, and `Safelist.isSafeAttribute()` is side effect free. [#&#8203;2475](https://github.com/jhy/jsoup/issues/2475) - `Cleaner` no longer duplicates enforced attributes when the input `Document` preserves attribute case. A case-variant source attribute is now replaced by the enforced attribute in the cleaned output. [#&#8203;2476](https://github.com/jhy/jsoup/issues/2476) - If a per-request SOCKS proxy is configured, jsoup now avoids using the JDK `HttpClient`, because the JDK would silently ignore that proxy and attempt to connect directly. Those requests now fall back to the legacy `HttpURLConnection` transport instead, which does support SOCKS. [#&#8203;2468](https://github.com/jhy/jsoup/issues/2468) - `Connection.Response.streamParser()` and `DataUtil.streamParser(Path, ...)` could fail on small inputs without a declared charset, if the initial 5 KB charset sniff fully consumed the input and closed it before the stream parse began. [#&#8203;2483](https://github.com/jhy/jsoup/issues/2483) - In XML mode, doctypes with an internal subset, such as `<!DOCTYPE root [<!ENTITY name "value">]>`, now round-trip correctly. The subset is preserved as raw text only; entities are not expanded and external DTDs are not loaded. [#&#8203;2486](https://github.com/jhy/jsoup/issues/2486) ##### Build Changes - Migrated the integration test server from Jetty to Netty, which actively maintains support for our minimum JDK target (8). [#&#8203;2491](https://github.com/jhy/jsoup/pull/2491) </details> --- ### Configuration 📅 **Schedule**: (UTC) - Branch creation - At any time (no schedule defined) - Automerge - At any time (no schedule defined) 🚦 **Automerge**: Disabled by config. Please merge this manually once you are satisfied. ♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox. 🔕 **Ignore**: Close this PR and you won't be reminded about this update again. --- - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box --- This PR has been generated by [Mend Renovate CLI](https://github.com/renovatebot/renovate). <!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4xMzIuMyIsInVwZGF0ZWRJblZlciI6IjQ0LjMuMiIsInRhcmdldEJyYW5jaCI6Im1haW4iLCJsYWJlbHMiOltdfQ==-->
renovate added 1 commit 2026-07-30 07:14:21 +02:00
Update dependency org.jsoup:jsoup to v1.23.1
continuous-integration/drone/pr Build is passing
a176e3a9b8
renovate force-pushed renovate/org.jsoup-jsoup-1.x from 9bf9e606e5 to a176e3a9b8 2026-07-30 07:14:21 +02:00 Compare
renovate changed title from Update dependency org.jsoup:jsoup to v1.22.2 to Update dependency org.jsoup:jsoup to v1.23.1 2026-07-30 07:14:24 +02:00
All checks were successful
continuous-integration/drone/pr Build is passing
You are not authorized to merge this pull request.
This pull request can be merged automatically.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin renovate/org.jsoup-jsoup-1.x:renovate/org.jsoup-jsoup-1.x
git checkout renovate/org.jsoup-jsoup-1.x
Sign in to join this conversation.