Skip to content

Latest commit

 

History

History
138 lines (97 loc) · 3.14 KB

File metadata and controls

138 lines (97 loc) · 3.14 KB

Getting Started

Installation

Install the CLI:

go install github.com/dotcommander/defuddle/cmd/defuddle@latest

Or add the library to your project:

go get github.com/dotcommander/defuddle

Quick Start

CLI

Extract the main content from any web page:

defuddle parse https://example.com/article

Convert to markdown:

defuddle parse https://example.com/article --markdown

Get structured JSON output with metadata:

defuddle parse https://example.com/article --json

Library

package main

import (
    "context"
    "fmt"
    "log"

    "github.com/dotcommander/defuddle"
)

func main() {
    result, err := defuddle.ParseFromURL(
        context.Background(),
        "https://example.com/article",
        nil,
    )
    if err != nil {
        log.Fatal(err)
    }

    fmt.Println(result.Title)
    fmt.Println(result.Content)
}

What You Get Back

Every parse returns a Result containing:

  • Content -- clean HTML with ads, navigation, and clutter removed
  • Title, Author, Published -- extracted from meta tags, Schema.org, and page structure
  • Domain, Favicon, Image -- site identity and social sharing image
  • WordCount -- CJK-aware word count of the extracted content
  • SchemaOrgData -- parsed JSON-LD structured data, when present
  • ParseTime -- extraction time in milliseconds

Core Concepts

Automatic Content Detection

Defuddle first honors a matching content selector or extraction bypass, then tries a supported site extractor. Other pages use Trafilatura v2.2.6 with native fallback to select readable article content. Defuddle restores supported rich formatting, resolves links, sanitizes HTML, and optionally converts it to Markdown. See configuration for processor gates and deprecated removal controls.

Site-Specific Extractors

For major platforms -- YouTube, Reddit, GitHub, ChatGPT, Claude, and others -- Defuddle uses purpose-built extractors that understand each site's DOM structure. When a URL matches a known platform, the site extractor runs instead of the general-purpose algorithm.

List all supported extractors:

defuddle extractors

Check which extractor matches a URL:

defuddle extractors --match https://www.youtube.com/watch?v=dQw4w9WgXcQ

Markdown Conversion

Request markdown output to get clean, readable text suitable for LLMs, note-taking, or further processing:

result, err := defuddle.ParseFromURL(ctx, url, &defuddle.Options{
    Markdown: true,
})
fmt.Println(*result.ContentMarkdown)
defuddle parse https://example.com --markdown

Batch Processing

Parse multiple URLs concurrently:

echo -e "https://example.com/a\nhttps://example.com/b" | defuddle batch
results := defuddle.ParseFromURLs(ctx, urls, &defuddle.Options{
    MaxConcurrency: 10,
})

Next Steps