Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .changeset/keep-heading-formatting.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
---
"@neuledge/context": patch
---

Keep bold, italic, link and inline-code text in section titles. Only a heading's top-level text was read, so words inside formatting or links were dropped from the title, which has the highest search weight. In the Python docs, 300 of 6,475 sections lost words (`Numeric Types — int, float, complex` became `Numeric Types — , , `), including 152 FAQ and guide sections whose heading is a link and which were titled "Introduction". Markdown was hit too: `## Using [superjson](...)` became "Using ". Heading permalinks such as Sphinx's "¶" stay out of the title.
23 changes: 23 additions & 0 deletions packages/context/src/build.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -76,6 +76,29 @@ Use brackets for dynamic segments.
expect(result.sections[2].sectionTitle).toBe("Dynamic Routes");
});

it("keeps formatted and linked text in section titles", () => {
const source = `## Using [superjson](https://github.com/blitz-js/superjson)

Serializes dates and maps.

## **Should I use \`generate\` or \`push\`?**

They are two different commands.

## The *strict* option

Turns on strict checks.
`;

const result = parseMarkdown(source, "docs/faq.md");

expect(result.sections.map((s) => s.sectionTitle)).toEqual([
"Using superjson",
"Should I use generate or push?",
"The strict option",
]);
});

it("uses docTitle from frontmatter", () => {
const source = `---
title: My Guide
Expand Down
19 changes: 9 additions & 10 deletions packages/context/src/build.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@
* Parses markdown/MDX, AsciiDoc, and reStructuredText files and chunks them by section.
*/

import type { Content, Heading, Root, Yaml } from "mdast";
import type { Content, Heading, PhrasingContent, Root, Yaml } from "mdast";
import remarkFrontmatter from "remark-frontmatter";
import remarkParse from "remark-parse";
import { unified } from "unified";
Expand Down Expand Up @@ -157,15 +157,14 @@ function extractFrontmatter(tree: Root): DocFrontmatter {

/** Get heading text from AST node. */
function getHeadingText(node: Heading): string {
let text = "";
for (const child of node.children) {
if (child.type === "text") {
text += child.value;
} else if (child.type === "inlineCode") {
text += child.value;
}
}
return text;
return node.children.map(getInlineText).join("");
}

/** Text of an inline node, including text nested in emphasis, strong and links. */
function getInlineText(node: PhrasingContent): string {
if (node.type === "text" || node.type === "inlineCode") return node.value;
if ("children" in node) return node.children.map(getInlineText).join("");
return "";
}

/** Convert AST nodes back to markdown text (simplified). */
Expand Down
28 changes: 28 additions & 0 deletions packages/context/src/html.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -35,3 +35,31 @@ describe("DocBook HTML examples", () => {
expect(parsed.sections[0]?.content).toContain("```sh\necho hello\n```");
});
});

describe("HTML headings", () => {
it("keeps linked heading text and drops permalink anchors", () => {
const parsed = parseHtml(
`<h1>Design FAQ</h1>
<h2><a class="toc-backref" href="#id3" role="doc-backlink">Why are Python strings immutable?</a><a class="headerlink" href="#why" title="Link to this heading">¶</a></h2>
<p>There are several advantages.</p>
<h3>Performance<a class="headerlink" href="#performance" title="Link to this heading">¶</a></h3>
<p>Strings of fixed size can be stored efficiently.</p>
<h2>Constants added by the <a class="reference internal" href="site.html#module-site"><code class="xref py py-mod docutils literal notranslate"><span class="pre">site</span></code></a> module<a class="headerlink" href="#constants" title="Link to this heading">¶</a></h2>
<p>The site module adds several constants.</p>
<h2 id="traits">Traits<a class="anchor" href="#traits">§</a></h2>
<p>Shared behaviour for types.</p>
<h2 id="setup"><span>Setup<a class="hash-link" href="#setup" aria-label="Direct link">&#8203;</a></span></h2>
<p>Install the package first.</p>`,
"faq/design.html",
);
expect(parsed.sections.map((s) => s.sectionTitle)).toEqual([
"Why are Python strings immutable?",
"Constants added by the site module",
"Traits",
"Setup",
]);
// Only section titles change: anchors in other headings stay in the content, which
// keeps its link ratio (and so the table-of-contents filter's verdict) as before.
expect(parsed.sections[0]?.content).toContain("[¶](#performance");
});
});
17 changes: 17 additions & 0 deletions packages/context/src/html.ts
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,23 @@ for (const tag of REMOVED_TAGS) {
turndown.remove(tag);
}

// Permalink anchors in section headings: "¶" (Sphinx, systemd), "§" (rustdoc), "#"
// (VuePress) or a zero-width space (Docusaurus). Section titles keep link text, so
// without this they end in the symbol. Only <h2>, which becomes the section title:
// anchors in other headings stay in the content as before. A rule, not remove():
// the link rule would match <a> first.
const PERMALINK_TEXT = /^[¶§#🔗]?$/u;
turndown.addRule("sectionPermalink", {
filter: (node) =>
node.nodeName === "A" &&
(node.getAttribute("href") ?? "").startsWith("#") &&
PERMALINK_TEXT.test(
(node.textContent ?? "").replace(/\u200b/g, "").trim(),
) &&
node.closest("h2") !== null,
replacement: () => "",
});

// DocBook emits bare <pre> elements; Turndown's code rule requires <pre><code>.
// Preserve their whitespace and prevent Markdown escaping of unit-file examples.
turndown.addRule("barePre", {
Expand Down
Loading