Skip to content

Latest commit

 

History

History

README.md

Xberg

html-to-markdown

High-performance HTML to Markdown converter for Swift via swift-bridge. Distributed through the Swift Package Manager; targets macOS 13+ and iOS 16+. Ship identical Markdown across every runtime while keeping idiomatic Swift APIs and Codable-friendly result types.

What This Package Provides

  • Same renderer as every binding — output matches Rust, Python, Node.js, Ruby, PHP, Go, Java, .NET, Elixir, R, Dart, Swift, Zig, C FFI, and WASM.
  • Structured conversion result — Markdown plus metadata, links, headings, images, tables, and warnings where the binding exposes them.
  • Production defaults — HTML is parsed with the Rust core, sanitized by default, and rendered without runtime-specific Markdown drift.
  • SwiftPM package — swift-bridge types for macOS and iOS codebases.

Installation

# SwiftPM URL consumption is not yet available (no published XCFramework).
# Build from source against the Rust static library:
git clone https://github.com/xberg-io/html-to-markdown.git
cd html-to-markdown
cargo build -p html-to-markdown-rs-swift --release
swift build --package-path packages/swift

Quick Start

Basic conversion:

import HtmlToMarkdown

let html = "<h1>Hello</h1><p>This is <strong>fast</strong>!</p>"
let result = try convert(html: html)
let markdown = result.content ?? ""
print(markdown)

With conversion options:

import HtmlToMarkdown

let options = try conversionOptionsFromJson(
    "{\"heading_style\":\"atx\",\"list_indent_width\":2,\"wrap\":true}"
)

let html = "<h1>Hello</h1><p>This is <strong>formatted</strong> content.</p>"
let result = try convert(html: html, options: options)
let markdown = result.content ?? ""
print(markdown)

Architecture

The converter routes each input through one of three tiers based on a fast prescan of the byte stream:

  1. Tier-1 — single-pass byte scanner. Handles 110+ HTML tags directly. Bails on any construct it cannot prove byte-equivalent to Tier-2.
  2. Tier-2 — DOM walker. Picks up Tier-1 bails and inputs the classifier rejected up front.
  3. Tier-3 — standards-conformant parser. Engaged for malformed HTML requiring full HTML5 repair.

The dispatcher is invisible to the caller. Output is byte-identical across tiers — enforced by a 116-snapshot oracle.

Capabilities

  • 16 languages, one Rust core. Rust, Python, Node.js, WASM, Java, Go, C#, PHP, Ruby, Elixir, R, Dart, Kotlin (Android), Swift, Zig, C ABI.
  • CommonMark-compatible Markdown with GFM-style tables.
  • Djot output: set output_format = "djot" (see Djot Output Format section below).
  • Real-HTML robust: unclosed tags, CDATA, custom elements, malformed entities, nested tables, mixed encodings handled without losing content.
  • Metadata extraction, visitor API, inline images, configurable preprocessing presets.
  • Per-group regression gates in CI: every PR runs the bench harness against per-group thresholds.

API Reference

Options

ConversionOptions – Key configuration fields:

  • heading_style: Heading format ("underlined" | "atx" | "atx_closed") — default: "atx"
  • list_indent_width: Spaces per indent level — default: 2
  • bullets: Bullet characters cycle — default: "-*+"
  • wrap: Enable text wrapping — default: false
  • wrap_width: Wrap at column — default: 80
  • code_language: Default fenced code block language — default: none
  • extract_metadata: Enable metadata extraction into result.metadata — default: true
  • output_format: Output markup format ("markdown" | "djot" | "plain") — default: "markdown"

Djot Output Format

The library supports converting HTML to Djot, a lightweight markup language similar to Markdown but with a different syntax for some elements. Set output_format to "djot" to use this format.

Syntax Differences

Element Markdown Djot
Strong **text** *text*
Emphasis *text* _text_
Strikethrough ~~text~~ {-text-}
Inserted/Added N/A {+text+}
Highlighted N/A {=text=}
Subscript N/A ~text~
Superscript N/A ^text^

Example Usage

Djot's extended syntax allows you to express more semantic meaning in lightweight text, making it useful for documents that require strikethrough, insertion tracking, or mathematical notation.

Plain Text Output

Set output_format to "plain" to strip all markup and return only visible text. This bypasses the Markdown conversion pipeline entirely for maximum speed.

Plain text mode is useful for search indexing, text extraction, and feeding content to LLMs.

Metadata Extraction

The metadata extraction feature enables comprehensive document analysis during conversion. Extract document properties, headers, links, images, and structured data in a single pass — all via the standard convert() function.

Use Cases:

  • SEO analysis – Extract title, description, Open Graph tags, Twitter cards
  • Table of contents generation – Build structured outlines from heading hierarchy
  • Content migration – Document all external links and resources
  • Accessibility audits – Check for images without alt text, empty links, invalid heading hierarchy
  • Link validation – Classify and validate anchor, internal, external, email, and phone links

Zero Overhead When Disabled: Metadata extraction adds negligible overhead and happens during the HTML parsing pass. Pass extract_metadata: true in ConversionOptions to enable it; the result is available at result.metadata.

Visitor Pattern

The visitor pattern enables custom HTML→Markdown conversion logic by providing callbacks for specific HTML elements during traversal. Pass a visitor as the third argument to convert().

Use Cases:

  • Custom Markdown dialects – Convert to Obsidian, Notion, or other flavors
  • Content filtering – Remove tracking pixels, ads, or unwanted elements
  • URL rewriting – Rewrite CDN URLs, add query parameters, validate links
  • Accessibility validation – Check alt text, heading hierarchy, link text
  • Analytics – Track element usage, link destinations, image sources

Supported Visitor Methods: 40+ callbacks for text, inline elements, links, images, headings, lists, blocks, and tables.

Links

Part of Xberg.io

  • Xberg — the open-source content-intelligence engine: text, tables, and metadata from 101 formats (115 file extensions), with OCR, transcription, and code intelligence. MIT.
  • Xberg Pro — a complete self-hosted content-intelligence backend in a single container. Commercial.
  • Xberg Enterprise — the distributed, governed content-intelligence platform, scaled on Kubernetes with team governance and support. Commercial.
  • crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
  • html-to-markdown — fast, lossless HTML→Markdown engine.
  • liter-llm — universal LLM API client with native bindings for 14 languages and 165 providers.
  • tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
  • alef — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.

Contributing

We welcome contributions! Please see our Contributing Guide for details on:

  • Setting up the development environment
  • Running tests locally
  • Submitting pull requests
  • Reporting issues

All contributions must follow our code quality standards (enforced via pre-commit hooks):

  • Proper test coverage (Rust 95%+, language bindings 80%+)
  • Formatting and linting checks
  • Documentation for public APIs

License

MIT License – see LICENSE. Copyright © Kreuzberg, Inc.

Support

If you find this library useful, consider sponsoring the project.

Have questions or run into issues? We're here to help: