Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .dockerignore
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,14 @@ web/.vite/
.env
.env.local
.cache/
logs/

# Local manuscript / screenshot test artifacts
novel.txt
novel.md
novel.docx
image*.png
image*.jpg

# Docs / tests are kept inside the image only when needed
docs/
Expand Down
51 changes: 51 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,57 @@ docker compose down -v # 同时删 cache 卷
- 后端镜像基于 `python:3.12-slim`,multi-stage 构建 + 非 root 用户运行,含 `/health` 探针。
- 前端镜像基于 `nginx:1.27-alpine`,serve 静态 `dist/` + 反代 `/api/`,SPA fallback 已配。
- LLM 响应缓存挂在 named volume `backend-cache`,跨容器重启保留。
- nginx `proxy_read_timeout = 300s`(覆盖 backend 的 LLM `read=180s` + 余量),长 beats / refine 调用不会被代理 504 截断。

#### Docker 部署验证清单

启动后按这个清单跑一遍,发现问题就能立刻定位到哪一层(容器 / 网络 / LLM):

```bash
# 1) 容器健康
docker compose ps
# 期望:backend 状态 (healthy),web 状态 running

# 2) 后端 /health(容器内)
docker compose exec backend python -c "import urllib.request as u; print(u.urlopen('http://127.0.0.1:8000/health').read())"
# 期望:b'{"status":"ok","version":"0.1.0"}'

# 3) /health 不在 /api 前缀下且 backend 8000 端口不对宿主开放,所以无法从宿主直接 curl;
# 第 2 步的容器内验证已足够确认存活。

# 4) 跑一次最小 /adapt(无 LLM key 也能通过规则路径)
curl -s -X POST http://localhost:8080/api/v1/adapt \
-H "Content-Type: application/json" \
-d '{"text":"第一章\n张三离开了小镇。\n\n第二章\n李四送他到车站。\n\n第三章\n两人挥手告别。","source_format":"txt","title":"测试","mode":{"fidelity":"medium","expansion":"balanced"}}' \
| head -c 200
# 期望:JSON 开头有 screenplay_yaml 字段

# 5) export 端点
curl -s -X POST http://localhost:8080/api/v1/export \
-H "Content-Type: application/json" \
-d '{"screenplay_yaml":"...上一步的 screenplay_yaml 字段值...","format":"fountain"}'
# 期望:JSON 含 content + mime_type=text/x-fountain

# 6) 缓存卷
docker volume inspect story2script_backend-cache
# 期望:Mountpoint 存在,第二次调相同 /adapt 应该几乎 0 ms(log 出 cache=hit)

# 7) 前端落地
curl -sI http://localhost:8080/ | head -3
# 期望:HTTP/1.1 200 OK + Cache-Control: no-cache(index.html)

# 8) 静态资源长缓存
curl -sI http://localhost:8080/assets/index-Df6cOw5V.css | grep -i cache
# 期望:Cache-Control: public, immutable + expires 1 年后
```

如果 LLM key 已配,再加一步:

```bash
# 9) 进容器看 LLM 日志
docker compose exec backend tail -n 20 logs/story2script.log
# 期望:每 stage 一行 llm.chat 或 pipeline.stage_ok,无 stage_reject 集中爆发
```

### 启用 LLM 增强(可选)

Expand Down
11 changes: 10 additions & 1 deletion docker/nginx.conf
Original file line number Diff line number Diff line change
Expand Up @@ -9,14 +9,23 @@ server {
# In docker-compose the upstream hostname is the service name "backend";
# outside compose set the Docker network DNS or override with an env-var
# rendered config.
#
# proxy_read_timeout must be ≥ the backend's LLM read timeout
# (default 180s, see src/story2script/llm/transport.py). Otherwise
# nginx kills long beats/refine calls with 504 before the model
# finishes streaming, which the frontend would surface as a generic
# network error and the backend would log as a phantom transport
# failure. 300s gives 120s headroom for the slowest legitimate case.
location /api/ {
proxy_pass http://backend:8000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_read_timeout 60s;
proxy_read_timeout 300s;
proxy_send_timeout 300s;
proxy_connect_timeout 10s;
}

# SPA fallback: every other route hits index.html so react-router
Expand Down
159 changes: 158 additions & 1 deletion docs/YAML_SCHEMA.md
Original file line number Diff line number Diff line change
Expand Up @@ -326,7 +326,164 @@ YAML 1.1 的二义性陷阱在第 7 节通过 `safe_dump` + round-trip 校验消
| ≥3 章节小说文本 | `metadata.chapter_count`;< 3 由 `FEW_CHAPTERS` 警告提示而非拒绝 |
| 自动转换为结构化剧本 | `scenes[].beats[]` 给出场景级与节拍级双层结构;每条带 kind 枚举 |
| YAML 格式 | 顶层即是 YAML;序列化由 §7 四道守则保障 |
| 让作者可编辑 | YAML 块状人类友好;`POST /validate` 接收作者改后的 YAML 做字段路径化校验 |
| 让作者可编辑 | YAML 块状人类友好;`POST /validate` 接收作者改后的 YAML 做字段路径化校验;`/editor` 提供 inline + scene-level 双轨编辑 |
| 可进一步打磨 | `source_type` + `confidence` + `needs_review` 三联:作者一眼看到哪里需要审,哪里可信 |
| 导出供后续创作 | `POST /api/v1/export` 把 Schema YAML 渲染为 Fountain / 纯文本 / Markdown,PDF 由前端浏览器打印生成 |

每一条赛题要求都有 Schema 字段或字段族对应,**Schema 不只是输出契约,也是赛题的"打分项映射表"**。

---

## 11. `quality_metrics` 响应包络

`POST /api/v1/adapt` 与 `/api/v1/adapt/upload` 不直接返回纯 YAML——它们包了一层响应对象,承载评分维度:

```jsonc
{
"screenplay_yaml": "schema_version: \"0.1.0\"\n...",
"schema_validation": { "valid": true, "errors": [] },
"quality_metrics": {
"chapter_count": 5,
"scene_count": 12,
"character_count": 4,
"detection_method": "llm",
"mean_chapter_confidence": 0.88,
"ai_stages_applied": ["chapters", "scenes", "characters", "beats", "refine"],
"llm_provider": "deepseek"
},
"warnings": [ ... ] // 同顶层 warnings[]
}
```

| 字段 | 类型 | 说明 |
|---|---|---|
| `chapter_count` / `scene_count` / `character_count` | int | 三个粒度的数量,前端徽标直接展示 |
| `detection_method` | enum | 章节检测路径:`llm` / `headings` / `regex` / `fallback` |
| `mean_chapter_confidence` | float | 章节 confidence 均值;< 0.5 时触发 `LOW_CONFIDENCE_DETECTION` |
| `ai_stages_applied` | string[] | 这次跑通 AI 路径的 stage 子集,`["chapters","scenes","characters","beats","refine"]` 的任意子集(最长 5) |
| `llm_provider` | string \| null | 请求时配置的 provider 名(`"deepseek"` / `"mimo"` / `null`) |

前端 `/app` 顶部的 `AI · deepseek · 5/5` 徽标就是 `llm_provider + len(ai_stages_applied) + 5` 的组合。这两个字段是 PR #35 引入的,schema 锚在 [`src/story2script/app/api/v1/models.py:QualityMetrics`](../src/story2script/app/api/v1/models.py)。

**为什么这两个字段不放进 YAML 顶层**:YAML 是"剧本数据",请求级 metrics 是"这次生成的元信息"——分开能让作者编辑 YAML 时不用担心碰到运行时元数据;也能让 `/validate` 校验的对象与 LLM 生成的对象保持一致(YAML round-trip 进 `/validate` 不会因为缺少 metrics 而失败)。

---

## 12. 编辑器:YAML 的就地修改 + 数据回流

`/editor` 路由把 Schema YAML 解析成富 UI,对它的每一次编辑都通过 [`web/src/lib/screenplay-patch.ts`](../web/src/lib/screenplay-patch.ts) 走一次 **js-yaml round-trip**:

```ts
const patched = patchScreenplayYaml(yaml, {
sceneIndex, beatIndex,
text: newText,
speaker: newSpeakerId,
source_type: "generated",
needs_review: false,
});
sessionStorage.setItem("story2script:draft", patched.yaml);
```

为什么不用字符串正则替换:beat 文本里出现正则元字符(`(冷冷地)` 这种 parenthetical)或者两条 beat 文本相同时,正则替换会静默改坏 YAML。round-trip 保证:

1. **结构正确**:永远是合法 YAML;
2. **可校验**:可立刻喂给 `POST /api/v1/validate` 验 schema;
3. **可恢复**:失败时返回 `null`,保持原 YAML 不变。

### `/app` ↔ `/editor` 数据回流

PR #61 修了一个隐藏的数据双源 bug:

- Editor 每次按键改 `sessionStorage["story2script:draft"]`;
- Adapt 的 result 面板 React state 里有一份 `result.screenplay_yaml`——LLM 第一次吐出来时的版本;
- 不同步的话,**导出剧本会导出改之前的版本**。

修复策略:Adapt 在 mount + tab focus + visibilitychange 三个时机各比对一次草稿与 result 的 YAML;不同就调 `/validate` 拿新的 `schema_validation`,setResult。`quality_metrics` **故意不刷**——它描述的是 LLM 怎么生成的草稿(章节数、AI stage 应用情况、置信度均值),与作者手动改无关。

---

## 13. 导出格式映射

`POST /api/v1/export` 把 Schema YAML 渲染成给演员 / 导演看的剧本格式。三种文本格式服务端生成,PDF 由前端浏览器打印渲染(避免捆绑 CJK 字体)。

### Schema 字段 → 各格式映射

| Schema 字段 | Fountain (`.fountain`) | 纯文本 (`.txt`) | Markdown (`.md`) |
|---|---|---|---|
| `metadata.title` | 标题页 `Title:` key | 居中显示在顶部 | `# <title>` |
| `metadata.mode` | 标题页 `Notes:` 备注 | 居中显示在标题下 | `> **模式**:` blockquote |
| `characters[]` | `/* DRAMATIS PERSONAE */` boneyard 注释 | `【角色表】` 段 | `## 角色表` + 项目列表 |
| `scenes[].chapter_title`(变更时)| `= 第一章` synopsis | `=` 横线 + 居中标题 | `## 第一章` |
| `scenes[].location` + `time` | `INT. <LOC> - <TIME>` SLUG(若首条 beat 没有则合成) | 同 | `### INT. <LOC> - <TIME>` |
| `scenes[].summary` | `= <summary>` synopsis | `〔<summary>〕` | `*<summary>*` 斜体 |
| `beats[].kind == "stage_direction"` | 普通段落(首条若 SLUG 形式则成为场景头)| 普通段落 | 普通段落 |
| `beats[].kind == "action"` | 普通段落(Fountain action 默认)| flush-left | 普通段落 |
| `beats[].kind == "dialogue"` | CHARACTER 大写 cue + (parenthetical) + body | 居中 cue + 缩进 body + wrap 60 列 | `**CHARACTER**` + `> *(parenthetical)*` + `> body` |
| `beats[].kind == "transition"` | `> CUT TO:` 强制标准化 | 右对齐 + 冒号 | `**CUT TO:**` |
| `beats[].kind == "narration"` | `NARRATOR (V.O.)` cue | `〈画外音〉` 前缀 | `> **【画外音】**` blockquote |
| `beats[].speaker == null` 的 dialogue | 退化为 `CHAR_<raw_id>` 大写(暴露 broken id 便于排查)| 同 | 同 |
| `metadata.generated_at` / `chapter_count` / `trace_id` | 不渲染(纯生成时元数据)| 不渲染 | 不渲染 |
| `source_type` / `confidence` / `needs_review` | 不渲染(剧本读者不关心 LLM 元数据)| 不渲染 | 不渲染 |

### Parenthetical 解析

Pipeline 在 dialogue text 里把情绪 / 动作提示放成 `(冷冷地) 你回来了?` 的前缀形式。Fountain 与 Markdown 渲染器都识别正则 `^[((]([^))]+)[))]` 把它拆成独立的 parenthetical 行。**纯括号 + 无对白主体**的退化情况(`(沉默)`)会保留在原 text,不拆——避免渲染出空白对白行。

### SLUG line 来源优先级

每个场景导出时,渲染器按以下顺序找 scene heading:

1. **首条 `stage_direction` beat 已是 `INT./EXT.` 开头** → 直接用它(pipeline 的 `_ensure_slug_line` 自动 prepend 的 SLUG 会落在这里,`source_type: inferred` `confidence: 0.7`);
2. **场景有 `location` 字段** → 合成 `INT. <LOCATION> - <TIME>`;
3. **都没有** → 跳过 SLUG,直接渲染场景正文(Fountain 解析器会把后续段落视作 action)。

### 422 schema 校验前置

`/export` 先跑 `_validate_against_schema`,不通过直接 422 + `detail.errors`。**渲染器不做 schema 容错**——如果 schema 错误能跨过校验,那一定是 schema 没覆盖到、应该补 schema 而不是补 renderer。

---

## 14. 字段交叉约束清单

下面这些约束在 schema 里以 `allOf + if/then`、`pattern`、`enum` 等组合方式落地。摘录到一处方便核对:

| 约束 | 触发位置 |
|---|---|
| `kind == "dialogue"` ⇒ `speaker` + `text` 必填 | beat 级 `if/then` |
| `speaker` 值必须匹配某个 `characters[].id` | 校验器附加检查(schema 本身允许任意 string,但运行期校验) |
| `first_appearance` 值必须匹配某个 `scenes[].id` | 同上 |
| `chapter_index` 必须 < `metadata.chapter_count` | 隐式:assemble 时按序号生成 |
| `confidence < 0.5` ⇒ `needs_review = true`(写入时自动)| Pipeline 写入约定,Schema 不强制(保留作者手动覆盖空间)|
| `aliases` 不含 `name` 本身 | Pipeline 写入时去重 |
| 每个 `scene.id` 必须形如 `scene_<int>` | `pattern: "^scene_\\d+$"` |
| 每个 `character.id` 必须形如 `char_<id>` 且全局唯一 | `pattern: "^char_[a-z0-9_]+$"` + 校验器查重 |
| `metadata.mode.fidelity` ∈ `{high, medium, low}` | enum |
| `metadata.mode.expansion` ∈ `{minimal, balanced, rich}` | enum |
| `metadata.source_format` ∈ `{txt, markdown, docx}` | enum |
| `source_type` ∈ `{original, inferred, generated, merged}` | enum |
| `kind` ∈ `{dialogue, action, stage_direction, narration, transition}` | enum |
| `confidence` ∈ `[0, 1]` | `minimum`/`maximum` |
| 所有对象 `additionalProperties: false` | schema 顶层约束,防止字段拼写漂移 |

### 校验器在哪

- **静态层(JSON Schema Draft 2020-12)**:所有数据类型、`required`、`pattern`、`enum` 都在这里。[`schemas/screenplay.schema.json`](../schemas/screenplay.schema.json)
- **动态层(业务校验)**:`speaker → character.id` 必须真存在、`character.id` 全局唯一这类**跨字段约束**,schema 表达成本太高,放在 `_validate_against_schema` 之外的 pipeline assemble 步骤。
- **`tests/test_schema_is_valid.py`** 锁住三件事:schema 自身是合法 Draft 2020-12、`examples/sample_output.yaml` 通过 schema、`/api/v1/adapt` 输出通过 schema。

---

## 15. 未来字段路线图

只列已设计但暂未实现的字段,方便 v0.2 引入时不破坏现有契约:

| 计划字段 | 位置 | 用途 |
|---|---|---|
| `scenes[].dramatic_unit` | scene 级 | "起承转合"标签,便于结构编辑 |
| `scenes[].emotional_tone` | scene 级 | 情绪基调(紧张 / 抒情 / 滑稽)供前端着色 |
| `characters[].relationships` | character 级 | 角色关系图,配对前端可视化 |
| `metadata.target_audience` | 文档级 | 影响 LLM prompt 风格 |
| `beats[].camera_hint` | beat 级 | 镜头提示(CU / WS / Dolly),保留给后续模型 |
| `metadata.duration_estimate` | 文档级 | 预估时长(基于 beats[] 估算)|

新增字段策略:**在 schema 里加 optional 字段**,旧客户端忽略不报错。`additionalProperties: false` 不必放宽——schema 列举的就是合法的字段全集。
Loading