第 07 / 11 章

扩展:把 rehypeCodeGroups 跑一遍

从一份最简 MDX 出发,摊开进入插件时的 HAST 全树,再按 i、k、out 把 rehypeCodeGroups 逐步执行完。不含 Prism。

⏱ 10 分钟2865 字

上一章用状态机把 rehypeCodeGroups 讲过一遍。若仍觉得「知道在合并,但树在脑子里对不上」,本章只做一件事:拿真实编译结果,对着源码逐步跑。

这是主线之外的扩展篇。读完应能自己在纸上模拟:给定一份 MDX,tree.children 每一步怎么变。rehypePrism 不在这里。

下面的树来自与站点相同的前半段管线(remark-gfm → rehype-slug),在即将进入 rehypeCodeGroups 时截取。解析器还会给每个节点挂 position(行列号),插件完全不用它,文中一律省略。

插件接到的是什么§

lib/mdx.tsx 里写的是:

rehypePlugins: [
  [rehypeSlug, { prefix: "" }],
  rehypeCodeGroups,  // ← 工厂本身,不是工厂的返回值
  rehypePrism,
]

unified 会先调用 rehypeCodeGroups(),得到真正改树的函数,再把整棵 HAST 喂进去:

export function rehypeCodeGroups() {
  return function (tree: Root) {
    const walk = (parent) => { /* … */ };
    walk(tree);
  };
}

所以「给到此函数」的 tree,类型是 HAST 的 Root:{ type: "root", children: [...] }。它已经不是 Markdown 字符串,也还不是 React 元素——只是一棵描述 HTML 的对象树。

两个会反复出现的节点形状:

type含义插件怎么认
"element"一个标签,如 p / pre / codeisElement:有 tagName
"text"纯文本,含块与块之间的 "\n"isBlankText:整段都是空白

围栏在这棵树上的固定形态是:

pre
 └─ code.language-xxx
      ├─ children: [ { type: "text", value: "源码\n" } ]
      └─ data.meta: "file=\"hi.js\""   ← 语言后面那串,还没解析

group / file 不在 pre 上,而在内层 code.data.meta。这是 remark/mdx 的约定,所以函数里总是先 node.children.find(isElement) 再 metaOf。

样本一:最简 MDX(没有 group)§

作者文件(可当成一篇只有正文的稿):

你好。

```js file="hi.js"
console.log(1)
```

进入 rehypeCodeGroups 时,完整 tree 如下(已去掉 position):

{
  "type": "root",
  "children": [
    {
      "type": "element",
      "tagName": "p",
      "properties": {},
      "children": [{ "type": "text", "value": "你好。" }]
    },
    { "type": "text", "value": "\n" },
    {
      "type": "element",
      "tagName": "pre",
      "properties": {},
      "children": [
        {
          "type": "element",
          "tagName": "code",
          "properties": { "className": ["language-js"] },
          "children": [{ "type": "text", "value": "console.log(1)\n" }],
          "data": { "meta": "file=\"hi.js\"" }
        }
      ]
    }
  ]
}

先把 root.children 看成一张有下标的表——process 只看这一层:

下标 i节点一句话
0<p>你好。</p>普通段落
1文本 "\n"块之间的换行,不是空白「段落」,但是空白文本
2<pre><code class="language-js">…</code></pre>唯一的围栏

函数不会先把整棵树「理解成文章」。它只是从 i = 0 走到末尾。

调用关系:三层盒子§

把源码叠成三层,后面逐步跑时只在最内层打转:

rehypeCodeGroups()                工厂,站点启动编译时调用一次
  └─ function (tree)              transformer,每篇文章一次
       └─ walk(parent)            先 process 当前层 children,再对子元素递归
            └─ process(nodes)     真正的线性扫描:一边读 nodes,一边写 out

walk 的两行:

parent.children = process(parent.children);
for (const c of parent.children) {
  if (isElement(c)) walk(c);
}

对样本一:先 walk(root) → process([p, \n, pre]) 得到新的 root.children;再进入 <p> 和 <pre> 各走一遍。<p> 里只有文本,process 原样送出;<pre> 里是 <code>,不是 <pre>,同样原样送出。合并只发生在「同一层里相邻的 <pre>」,递归是为了列表、引用块里的围栏也能被扫到。

下面只盯第一次 process(root.children)。

样本一逐步执行§

开始:nodes 长度 3,out = [],i = 0。

i = 0:段落,原样搬§

node 是 <p>。isPre(node) 为假(tagName !== "pre")。

if (!isPre(node)) {
  out.push(node);
  i += 1;
  continue;
}

结果:out = [p],i = 1。对象没被复制,只是引用推进了新数组。

i = 1:换行文本,同样原样搬§

node 是 { type: "text", value: "\n" },不是 element,更不是 pre。同一分支:out = [p, \n],i = 2。

i = 2:碰到 <pre>,进入围栏分支§

isPre(node) 为真。下一句:

const code0 = node.children.find(isElement);

pre.children 只有一个元素,就是那颗 <code>。code0 找到了。

若找不到(空的 <pre>),会原样 push 后 continue——防御畸形树。样本一用不到。

读 meta、拼 file0§

const meta0 = metaOf(code0);
const file0 = { name: meta0.file, lang: langOf(code0), code: codeText(code0) };

metaOf 读 code0.data.meta,得到字符串 file="hi.js",交给 parseMeta:

  1. 正则切出 token:file="hi.js"
  2. 含 =,key 为 file,value 去掉引号得 hi.js
  3. 没有 group / preview / title

于是:

{
  "meta0": { "group": null, "file": "hi.js", "preview": false, "title": null },
  "file0": { "name": "hi.js", "lang": "js", "code": "console.log(1)\n" }
}

langOf 看 className 里以 language- 开头的项,切掉前缀并小写 → "js"。
codeText 递归拼文本节点 → "console.log(1)\n"(围栏末行换行保留在原文里)。

没有 group:只装饰,不向后看§

if (!meta0.group) {
  decoratePre(node, node.children, [file0], { group: null, preview: meta0.preview });
  out.push(node);
  i += 1;
  continue;
}

decoratePre 做两件事:

  1. 在 pre.properties 上写入 data-files(JSON 字符串)。group 为 null、preview 为 false,所以不写 data-group、data-preview。
  2. pre.children = codes。这里传入的是原来的 node.children,所以内层结构不变,仍是一个 <code>。

out 变成 [p, \n, 已装饰的 pre],i = 3。循环结束。

process 返回 out,赋回 root.children。树还是三个顶层节点,但第三个 <pre> 多了属性:

{
  "type": "root",
  "children": [
    { "type": "element", "tagName": "p", "…": "你好。" },
    { "type": "text", "value": "\n" },
    {
      "type": "element",
      "tagName": "pre",
      "properties": {
        "data-files": "[{\"name\":\"hi.js\",\"lang\":\"js\",\"code\":\"console.log(1)\\n\"}]"
      },
      "children": [
        {
          "type": "element",
          "tagName": "code",
          "properties": { "className": ["language-js"] },
          "children": [{ "type": "text", "value": "console.log(1)\n" }],
          "data": { "meta": "file=\"hi.js\"" }
        }
      ]
    }
  ]
}

要点:没有 group 不是「什么都不做」。单文件也必须带 data-files,后面的 CodeBlock 才能用同一套解析复制源码。子节点尚未高亮——那是下一个插件的事。

单文件路径不会 delete child.data。meta 仍挂在 code 上。有 group 的合并路径才会清掉(见下一节)。对客户端几乎无影响:React 不会把 data 字段渲染成 DOM 属性。

样本二:相邻同组(真正合并)§

作者:

先看一段说明。

```js group="demo" file="a.js"
const a = 1
```

```css group="demo" file="b.css"
.a { color: red }
```

进入插件时的 root.children:

下标节点code.data.meta
0<p>先看一段说明。</p>—
1文本 "\n"—
2<pre> / language-jsgroup="demo" file="a.js"
3文本 "\n"—
4<pre> / language-cssgroup="demo" file="b.css"

完整树:

{
  "type": "root",
  "children": [
    {
      "type": "element",
      "tagName": "p",
      "properties": {},
      "children": [{ "type": "text", "value": "先看一段说明。" }]
    },
    { "type": "text", "value": "\n" },
    {
      "type": "element",
      "tagName": "pre",
      "properties": {},
      "children": [
        {
          "type": "element",
          "tagName": "code",
          "properties": { "className": ["language-js"] },
          "children": [{ "type": "text", "value": "const a = 1\n" }],
          "data": { "meta": "group=\"demo\" file=\"a.js\"" }
        }
      ]
    },
    { "type": "text", "value": "\n" },
    {
      "type": "element",
      "tagName": "pre",
      "properties": {},
      "children": [
        {
          "type": "element",
          "tagName": "code",
          "properties": { "className": ["language-css"] },
          "children": [{ "type": "text", "value": ".a { color: red }\n" }],
          "data": { "meta": "group=\"demo\" file=\"b.css\"" }
        }
      ]
    }
  ]
}

i = 0、i = 1 与样本一相同:段落和换行进 out。关键从 i = 2 开始。

i = 2:第一个带 group 的 pre§

取出 code0、meta0、file0:

{
  "meta0": { "group": "demo", "file": "a.js", "preview": false, "title": null },
  "file0": { "name": "a.js", "lang": "js", "code": "const a = 1\n" }
}

meta0.group 为 "demo",不走单文件分支,进入吞并:

const group = meta0.group;          // "demo"
const pres = [node];                // 第一个 pre(下标 2)
const files = [file0];              // 已有 a.js
const codes = [];                   // 先空着,合并结束再装填
let preview = meta0.preview;        // false
let k = i + 1;                      // 从 3 往后看

注意三个数组职责不同:

数组装什么最后干什么
pres被吞掉的那些 <pre> 节点从里面抠 <code>
files{ name, lang, code } 原文清单JSON.stringify 进 data-files
codes实际的 <code> 元素成为第一个 <pre> 的新 children

k 循环:决定组有多长§

for (; k < nodes.length; k++) {
  const n = nodes[k];
  if (isBlankText(n)) continue;
  if (!isPre(n)) break;
  const c = n.children.find(isElement);
  if (!c) break;
  const m = metaOf(c);
  if (m.group !== group) break;
  pres.push(n);
  files.push({ name: m.file, lang: langOf(c), code: codeText(c) });
  if (m.preview) preview = true;
}

k = 3:nodes[3] 是 "\n"。isBlankText 为真(type === "text" 且整段空白)。continue——不 break,组还没结束。k 由 for 头加到 4。

这就是「相邻可以夹空白」的全部实现:空白被跳过,扫描继续。若这里是 <p>,会走到 !isPre(n) 然后 break,两个围栏就不会合并。

k = 4:nodes[4] 是第二个 <pre>。不是空白;是 pre;内有 <code>;metaOf 得到 group === "demo",与当前组相同。于是:

  • pres 变成 [pre-js, pre-css]
  • files 多一项 { name: "b.css", lang: "css", code: ".a { color: red }\n" }
  • 该块没有 preview,preview 仍为 false

k = 5:k < nodes.length 为假,循环结束。此时 k 停在 5(第一个「不能并入」的位置;这里恰好是数组尾)。

end:再吃掉组后面的空白§

let end = k;
while (end < nodes.length && isBlankText(nodes[end])) end += 1;

k 已是 5,没有更多节点,end = 5。若组后面还有 "\n",这里会把它们算进「已消费区间」,避免合并后 out 里留下一段多余空白。

装填 codes,只留第一个 pre§

for (const p of pres) {
  for (const child of p.children) {
    if (isElement(child)) {
      delete child.data;
      codes.push(child);
    }
  }
}
decoratePre(node, codes, files, { group, preview });
out.push(node);
i = end;

pres 里两个 pre 各有一个 <code>,codes 成为 [code-js, code-css]。各自的 data.meta 被删掉——契约已经写进 files,不必把 meta 再带到客户端。

decoratePre 作用在第一个 pre(i = 2 那个)上:

  • data-group="demo"
  • data-files 为两个文件的 JSON
  • children 换成两个 <code>

第二个 <pre> 不会 push 进 out。i = end = 5,外层 while 结束。

顶层从 5 个节点变成 3 个:

{
  "type": "root",
  "children": [
    { "type": "element", "tagName": "p", "…": "先看一段说明。" },
    { "type": "text", "value": "\n" },
    {
      "type": "element",
      "tagName": "pre",
      "properties": {
        "data-group": "demo",
        "data-files": "[{\"name\":\"a.js\",\"lang\":\"js\",\"code\":\"const a = 1\\n\"},{\"name\":\"b.css\",\"lang\":\"css\",\"code\":\".a { color: red }\\n\"}]"
      },
      "children": [
        {
          "type": "element",
          "tagName": "code",
          "properties": { "className": ["language-js"] },
          "children": [{ "type": "text", "value": "const a = 1\n" }]
        },
        {
          "type": "element",
          "tagName": "code",
          "properties": { "className": ["language-css"] },
          "children": [{ "type": "text", "value": ".a { color: red }\n" }]
        }
      ]
    }
  ]
}

之后 walk 进入这个 <pre>,process 看到的 children 是两个 <code>:都不是 pre,原样写回。树稳定。

客户端将看到一个 pre(映射成一个 CodeBlock):Tab 文案来自 data-files[].name,当前高亮 DOM 来自 children[activeIdx]。两个来源下标对齐,是因为装填 codes 与装填 files 用的是同一段 pres 顺序。

样本三:同名 group 被段落切开§

进入时:

下标节点
0<pre> group=a file=a.js
1"\n"
2<p>中间打断。</p>
3"\n"
4<pre> group=a file=b.js

i = 0 进入吞并,k = 1 跳过空白,k = 2 是 <p> → !isPre → break。pres 里只有第一块。各自装饰成独立的单组 pre(仍带 data-group="a",但 data-files 只有一项)。

group 不是全局钥匙,只是当前层、当前扫描窗口里的相等判断。

把 process 收成一张决策表§

对 nodes[i],从上到下只走一条:

判断动作i 怎么走
不是 <pre>out.pushi + 1
是 <pre> 但没有元素子节点out.pushi + 1
有 code,但 meta.group 为空decoratePre(单文件 data-files)后 pushi + 1
有 group用 k 吞相邻同组;收集 codes/files;装饰第一个 pre;丢掉其余 pre 与组后空白i = end

k 循环内部:

判断动作
空白文本跳过,继续
非 <pre>组结束
<pre> 无 code组结束
group 字符串不等(含 null)组结束
其余并入 pres / files,组内任一 preview 则整组可预览

为何不在原数组上 splice§

process 另建 out,最后一次性替换。若边扫边从 nodes 里删掉已吞的 <pre>,k 与 i 会互相踩。新数组只含「还要出现在这一层的节点」:段落、换行、以及每个组留下的那一个 <pre>。

和主线的关系§

本章把 rehypeCodeGroups 从「合并算法」落实到「这一次 process 的内存怎么变」。上一章仍负责双轨契约、parseMeta 全表、以及和 rehypePrism 的顺序。下一章回到主线:app/ 如何把散篇与系列章节接到同一条 MdxContent。