feat: 验证英语分词、原文定位与离线词典 (#3)

This commit is contained in:
ila
2026-09-10 20:59:18 +08:00
parent 96ac5eb618
commit b7c976eb75
21 changed files with 1020 additions and 12 deletions
+2 -1
View File
@@ -258,7 +258,7 @@ MVP 内所有单元任务通过后才能做 MVP 集成验收;MVP 通过后才
- 治理模式:轻量。数据库:MySQL 8(用户于 2026-09-10 确认);具体小版本在工程验证后锁定。
- 远端:https://git.ilapage.cn/OPC/lexgo.git;分支 main。不得把邻接 dev_harness 工作区当成本项目工作区。
- 工程基础 #2 已通过用户验收:server 基于指定 go-admin 选用模型扩展账号/会话 API,admin 复用 go-admin-ui,learner 为独立 Vue 3 + TypeScript + Vite 工程。默认英语;阅读、导入、词典、复习及 Python NLP 尚未实现或验证。
- 工程基础 #2 已通过用户验收:server 基于指定 go-admin 选用模型扩展账号/会话 API,admin 复用 go-admin-ui,learner 为独立 Vue 3 + TypeScript + Vite 工程。默认英语;阅读、导入、词典与复习尚未接入产品;#3 已完成独立 Python NLP/词典验证小样,待验收。
- 原四份研究保留为历史参考;PostgreSQL 建议被 MySQL 8 决策覆盖,U/A/N 索引用于追踪而不是批准所有范围。
- 用户/语言数据所有权、Unicode 原文位置、任务和复习幂等、完整备份恢复是后续方案的必要验收边界。
- 当前 MCP 连接其他 Gitea 站点,需使用目标站点 API 时记录原因;凭据仅从安全配置进入进程。
@@ -278,3 +278,4 @@ MVP 内所有单元任务通过后才能做 MVP 集成验收;MVP 通过后才
- 已验证 MySQL 8.4.3,本机 127.0.0.1:3308;开发库 lexgo_dev、测试库 lexgo_test_issue2。密码只从环境或忽略的 .env.local 读取。迁移测试只能使用 lexgo_test_ 前缀专用库,不能借用其他数据库。
- 后端命令使用 `python scripts/server.py migrate|bootstrap|serve|build|test|test-integration`;仅显式 migrate 修改表。bootstrap 只接受尚无账号的 LexGo 库,不覆盖已有管理员。Go 1.26.5、Node 22.22.1、pnpm 9.15.1;两端分别构建。
- #18 登录日志与操作审计已通过用户验收:schema v2 显式迁移;日志只保存白名单字段,禁止保存凭据、请求/响应正文及私人学习内容。仅管理员查询,默认保留 90 天;启动/每小时及 `python scripts/server.py audit-cleanup` 仅清理两张审计表的过期记录。
- #3 独立小样位于 `spikes/english/`,使用 `.local/nlp-venv/Scripts/python.exe`(3.12.12)运行;固定 spaCy 3.8.7、英语模型 3.8.0、NLTK 3.9.2、WordNet 3.0。资源仅显式准备时下载,摘要见 resources.json。不得把本机无账号的实验接口用于正式学习端;后续集成仍需 Go 授权、数据归属和任务设计。原文不归一化,位置区分 cp/UTF-8/UTF-16,lemma 不自动合并学习状态。
+1
View File
@@ -5,6 +5,7 @@
已确认:**DevHarness 轻量模式、MySQL 8、go-admin 管理端**。工程基础 #2 已通过验收:两端用户名登录、学习账号管理、可撤销会话和本人英语空空间。管理端基于指定 go-admin/go-admin-ui 选用模块,学习端为独立 Vue 3 + TypeScript + Vite 工程,共用 Go 后端和 MySQL 8.4.3。#18 登录日志与操作审计已通过用户验收,支持管理员查询和 90 天保留清理。阅读、导入、词典与复习尚未实现。MVP 定位为“支持多账号、数据独立的自托管学习工具”,先邀请少量用户使用;F01~F12 已确认,X 系列后置。
- [文档入口](docs/README.md) · [线上 Wiki](https://git.ilapage.cn/OPC/lexgo/wiki/Home)
- [英语分词与离线词典验证小样](spikes/english/README.md)(#3 待验收,独立本机入口)
- [项目档案](docs/00-project-profile.md) · [需求总览](docs/09-product-requirements-overview.md)
- [工作量估算](docs/10-workload-estimate.md):#2 验收后原范围剩余 46~75 人日,新增 #18 的 3~5 人日计划后为 49~80 人日;技术验证后重估,旧全量研究仅供参考。
- [四阶段实施总览 #16](https://git.ilapage.cn/OPC/lexgo/issues/16):14 张单元工单,工程基础 → 技术验证 → 首条学习闭环 → 补齐 MVP;原型 v1 已获用户验收。两端使用账号(用户名)+密码登录,不要求邮箱。
+22 -5
View File
@@ -2,8 +2,8 @@
generated: true (请先修改 Gitea Wiki,禁止直接编辑本文件)
wiki_page: Architecture-and-Code-Map
wiki_url: https://git.ilapage.cn/OPC/lexgo/wiki/Architecture-and-Code-Map.-
wiki_revision: 3b6f6db7f0fe1019b5811f70af8d6174cec0b950
synchronized_at: 2026-09-10T12:36:57Z
wiki_revision: 0806b7fa3ee8606020c9d7185e26c059c0b761c9
synchronized_at: 2026-09-10T12:58:33Z
<!-- gitea-wiki-mirror:end -->
# 架构与代码地图
@@ -23,7 +23,7 @@ Go 承担业务与后台任务,浏览器提供阅读学习界面,NLP 保留
| `dev_scripts/harness.py`、`dev_scripts/wiki_docs.py` | 原样复制的 DevHarness 工具 | 治理工具入口,无业务 API |
| `tests/` | 上游治理工具与文档结构测试 | 不验证阅读、NLP 或 SRS |
不存在产品入口、数据库迁移或前端页面。拟定职责:identity(身份)、library(书库)、ingestion(导入)、lexicon(全局词典)、vocabulary(个人词语)、review(复习)、progress(统计)、administration(管理)。底座固定后再决定具体目录。
工程入口与已实现模块见下方 #2/#18;学习领域拟定职责:identity(身份)、library(书库)、ingestion(导入)、lexicon(全局词典)、vocabulary(个人词语)、review(复习)、progress(统计)、administration(管理)。底座固定后再决定具体目录。
## 两条主要执行路径
@@ -47,7 +47,7 @@ Go 承担业务与后台任务,浏览器提供阅读学习界面,NLP 保留
## 学习端与管理端的目标架构
技术方向已纳入本轮方案,以下是目标结构,尚无对应产品目录或可运行应用。
以下是目标结构;账号、管理端与学习空空间已由 #2 实现。英语 NLP 已完成 #3 独立验证,尚未接入业务 API。
```mermaid
flowchart TD
@@ -56,7 +56,7 @@ flowchart TD
API --> B[独立学习业务模块]
B --> DB[(MySQL 8)]
B --> W[后台任务:具体选型待验证]
W --> N[NLP 服务:Python 方案待确认]
W --> N[NLP:Python 小样已验证,生产集成待实施]
```
| 交付部分 | 建设方式 | 复用与自建边界 |
@@ -131,3 +131,20 @@ server/app/lexgo/audit.go 定义两类白名单字段日志、筛选分页、失
database.go 显式迁移至 v2,两张新增表均以 created_at/id 建立排序清理索引,账号字段建查询索引,无业务表级联删除。cmd/lexgo/main.go 的服务进程在启动和每小时执行审计清理,每次最多运行一分钟、每批删除 1000 条,仅影响过期审计记录。
admin/src/views/AuditLogs.vue 通过 kind 复用登录/操作列表;audit-logs.mjs 负责筛选编码和请求序号,session.mjs 继续进行管理员及会话 generation 校验。切换页面/账号清空日志,普通翻页保留总数,防止分页组件跳回第一页。菜单与标题按当前路由显示。
## 英语分词与本地词典验证(#3,待验收)
`spikes/english/` 是独立可运行验证小样,不是学习端生产功能。推荐后续采用 Python 3.12.12、spaCy 3.8.7、en_core_web_sm 3.8.0(保留 tok2vec/tagger/attribute_ruler/lemmatizer,停用 parser/ner)和 NLTK 3.9.2 读取 WordNet 3.0。Go 继续管理用户、权限、任务和持久数据,后续通过显式契约调用 NLP;本单未新增 Go API、MySQL 表或常驻部署实例。
| 文件 | 作用 |
|---|---|
| engine.py | 原文分词、lemma、三个位置单位、直接/lemma 查词;只读本地资源 |
| app.py、index.html、app.mjs、view.mjs、style.css | loopback 临时 HTTP 小样、输入/阅读/查词结果;单进程串行,输入不落盘 |
| resources.json、setup_resources.py、requirements.lock | 固定版本、来源和 SHA256;显式联网准备,运行期无自动下载 |
| test_engine.py、test_app.py、view.test.mjs | 真实模型离线验证、HTTP 边界与浏览器偏移/迟到响应测试 |
| benchmark.py、benchmark-result.json | 虚构语料的候选对照、长文/查询实测及环境样本 |
WordNet 使用 ZIP 内原始 index/data/exception 文件,不使用 SysDict 或新增业务库。NLTK 默认 synsets 会隐式词形还原,本小样直接读取其固定版本索引以区分 exact 和显式 lemma;禁用依赖全局 corpus 的 OMW 跨版本映射,只接受 WordNet 3.0。升级 NLTK 或词典时必须重跑契约测试。
候选比较:正则分词+WordNet 默认名词 morphology 依赖少、速度快,但不具备上下文判断,缩写和词性歧义处理弱;纯 Go 规则同样需要自行维护这些语言规则。本次 spaCy 在 12 个明确样例中答对 11 个,基线 6 个,因此推荐保留独立 Python NLP 边界。样例量不足以证明总体准确率;不宣称部署或正式阅读功能已完成。
+21 -2
View File
@@ -2,8 +2,8 @@
generated: true (请先修改 Gitea Wiki,禁止直接编辑本文件)
wiki_page: Business-Rules-and-Glossary
wiki_url: https://git.ilapage.cn/OPC/lexgo/wiki/Business-Rules-and-Glossary.-
wiki_revision: 7061627fed9b5acd74fbd2369f27acadc0fd2e2a
synchronized_at: 2026-09-10T12:36:59Z
wiki_revision: 493544b5fd723c028d9173ecf590b66fea72c20a
synchronized_at: 2026-09-10T12:58:34Z
<!-- gitea-wiki-mirror:end -->
# 业务规则与术语
@@ -80,3 +80,22 @@ M0 固定首发语言语料、词条身份规则、短语选择与重叠规则
- 日志绝不保存密码、token、Cookie、请求/响应正文、错误堆栈或私人学习内容。合法账号、IP 属于本功能必要的审计数据,仅管理员可查询。
- GET /api/v1/login-logs 与 /operation-logs:未登录 401,学习者 403;page 默认 1,limit 默认 20、最大 100。username 精确匹配;操作日志匹配操作人或目标账号。result 为 success/failure,action 限定枚举,from/to 为 RFC3339。返回 data.items/total/page/limit,按时间和编号倒序,时间按毫秒存储、浏览器按本地时区展示。
- 固定保留最近 90 天,查询即排除过期记录;默认起止为保留边界和当前时间。清理只删除 created_at 严格早于边界的两表记录,不影响账号、空间、会话。无清空全部或导出按钮;本期不开放保留时长配置。
## #3 英语位置与查询实验契约 v1
实验版本 `english-spike-v1`。POST /analyze 接收 {text},返回 status、contract_version、original_text、text_sha256 和 tokens;仅为本机小样接口,不是生产 API。原文以收到的字符串为准,不先做 NFC、大小写、换行或空白归一化;SHA256 对原文 UTF-8 字节计算。tokens 连续覆盖全文,拼接 text 必须逐字符等于原文,空白也有独立区间。空串合法;最多 100000 Unicode code point,拒绝孤立代理项。kind 为 word/space/punctuation;word 是本小样的可点 token 类别,也可能包含数字或 emoji,不保证是自然语言词条。
每个 token 提供 text、lemma、kind 与半开区间 [start,end):
| 字段后缀 | 单位与使用方 |
|---|---|
| cp | Unicode code point,Python 字符串索引;不是用户感知字形 |
| utf8 | UTF-8 字节,可供 Go string 字节切片 |
| utf16 | UTF-16 code unit,JavaScript String.slice / DOM 文本位置 |
例:原文 `A🙂é`,emoji 的 cp=[1,2)、utf8=[1,5)、utf16=[1,3);后面的 e 加组合重音共两个 code point,cp=[2,4)、utf8=[5,8)、utf16=[3,5)。不能把这些单位混用,也不能把组合字符或 ZWJ 序列的 code point 数当作可见字符数。浏览器先逐 token 校验 UTF-16 切片并检查完整重建再展示;textarea 会按浏览器规范将换行转为 LF,所以 HTTP 契约保证收到的原文,不承诺还原剪贴板进入 textarea 前的 CRLF。服务端 CRLF 原文测试单独覆盖。
POST /lookup 接收 {surface,lemma?}。查词键单独 casefold/NFC/弯撇号转 ASCII,不改变原文位置;先精确查询 surface,再尝试调用方提供的 lemma。结果 status 为 exact、lemma、not_found 或 resource_missing,含 matched_form 和最多 12 条 entries(lemma/pos/definition/examples)。`dog` 直接命中;点击 `went` 可用上下文 lemma `go` 回退;手动只输入 `went` 不猜词性而返回未找到。词典缺失和模型缺失分别标识 wordnet/model,不能伪装成查无结果。
实验查询无用户状态、写入或缓存私人输入;重复查询确定性返回,原文哈希可检测文本版本变化,但尚未定义生产 token ID、任务幂等或个人词语合并规则。lemma 不等于学习状态身份,禁止自动合并原词/词元。WordNet 仅英英释义,按 n/v/a/r 与原生 sense 顺序截取,不做上下文义项排序、翻译或发音;`The leaves fell.` 的 leaves 实测被模型错误还原为 leave,此限制保留供后续用户选择/修正方案参考。
+35 -2
View File
@@ -2,8 +2,8 @@
generated: true (请先修改 Gitea Wiki,禁止直接编辑本文件)
wiki_page: Local-Development-and-Verification
wiki_url: https://git.ilapage.cn/OPC/lexgo/wiki/Local-Development-and-Verification.-
wiki_revision: 8857cf2545a3de6f1efa07d88b920a8301080e56
synchronized_at: 2026-09-10T12:43:25Z
wiki_revision: bf3483ec95bb73d522776d89dd2c20f9d117adb5
synchronized_at: 2026-09-10T12:58:36Z
<!-- gitea-wiki-mirror:end -->
# 本地开发与验证
@@ -192,3 +192,36 @@ supervisor 直接管理编译后的 Go 进程,运行时不调用 Python。数
#18 用户验收:2026-09-10T20:40:04+08:00 用户确认日志通过验收(工单评论 7591),包含此前待人工检查的交互。未重新运行自动化测试,未更改其历史结果,未合并 PR 或发布生产。
## #3 英语离线验证入口与复现
从仓库根目录执行(uv 与 Node 已安装,不能使用本机默认 Python 3.8):
```powershell
uv venv --python 3.12.12 .local/nlp-venv
uv pip install --python .local/nlp-venv/Scripts/python.exe -r spikes/english/requirements.lock
.local/nlp-venv/Scripts/python.exe spikes/english/setup_resources.py
uv pip install --python .local/nlp-venv/Scripts/python.exe --no-deps .local/nlp-resources/en_core_web_sm-3.8.0-py3-none-any.whl
.local/nlp-venv/Scripts/python.exe -m unittest discover -s spikes/english -v
node --test spikes/english/view.test.mjs
.local/nlp-venv/Scripts/python.exe spikes/english/benchmark.py
.local/nlp-venv/Scripts/python.exe spikes/english/app.py
```
打开 http://127.0.0.1:5183/,默认虚构样例,分析后点击 went 应出现 go 与“按原形查询”;dog 直接命中,zzzxqvfiction 未找到。`--resources .local/absent-resources` 可验证词典缺失,`--port` 可更换临时端口。模型缺失、词典缺失、非法输入及内部错误不输出路径/正文。仅 loopback,Host/Origin 校验,禁跨域、无缓存、无访问日志、连接读超时 10 秒。未配置 supervisor;停止该临时进程即可回退,既有服务和数据不变。
首次准备需要联网,失败可重跑;资源文件通过固定 SHA256 校验后使用。模型 3.8.0 MIT,WordNet 3.0 ZIP 完整保留 LICENSE/版权/免责声明,spaCy MIT、NLTK Apache-2.0。固定资源 URL 和摘要见 spikes/english/resources.json,原始许可与来源见该目录 README。词典为英英格式,不是中文翻译库。
2026-09-10 实测:11 项 Python 测试和 2 项 JavaScript 测试通过。真实模型/词典测试及 benchmark 禁止 socket connect,验证运行期无在线翻译依赖;不是整机断网测试。浏览器已实测展示分词、点击 went→go、手动查询无结果。外部 spaCy/Click 有一条 DeprecationWarning,未影响测试结果。Windows 10 19044,Intel Family 6 Model 140、8 逻辑核,Python 3.12.12;完整环境、UTC 时间和样本保存在 benchmark-result.json。
| 测量 | 本次样本 |
|---|---|
| 冷进程 Engine 加载(含 import,文件系统缓存可能已热) | 2355 ms |
| 首次分析 / 首次 dog 查询 | 见 benchmark-result.json(各 1 次) |
| 100000 code point(106095 UTF-8 字节),spaCy 3 次 | 中位 1433 ms,约 6.98 万 cp/s |
| 同文正则+WordNet morphology,3 次 | 中位约 336 ms;未包含三位置转换,非完全等价负载 |
| dog 查询,热进程 100 次 | 中位 0.0149 ms,p95 0.023 ms |
| 12 个显式 lemma 样例 | spaCy 11/12、基线 6/12;保留 leaves 错误 |
推荐 Python NLP,但该样本不代表一般准确率、生产并发能力或延迟保证。尚未验证正式 Go/Python 调用、长任务持久化、移动端划词(#4)、英汉词典及生产部署。
+5 -2
View File
@@ -2,8 +2,8 @@
generated: true (请先修改 Gitea Wiki,禁止直接编辑本文件)
wiki_page: Home
wiki_url: https://git.ilapage.cn/OPC/lexgo/wiki/Home
wiki_revision: d113bfc4d8d33feced5b84f098a5c50ba51df0ba
synchronized_at: 2026-09-10T12:43:18Z
wiki_revision: c5a40a06bd288faab9c0a3cbebb3bc2713bc13e7
synchronized_at: 2026-09-10T12:58:28Z
<!-- gitea-wiki-mirror:end -->
# LexGo 文档入口
@@ -57,3 +57,6 @@ Quant-UX 原型 v1 已通过用户验收。[桌面预览](https://qux.ilapage.cn
日志审计 #18 于 2026-09-10T20:40:04+08:00 获用户验收并关闭;#16 已更新完成索引。下一阶段 #3/#4 尚未开始。
英语分词/原文定位/本地词典 #3 已完成独立可运行小样,待用户验收。代码与复现命令位于 spikes/english;临时入口 http://127.0.0.1:5183/。推荐 Python spaCy 英语模型与 WordNet 3.0 的离线组合;尚未接入正式学习端,#4 划词验证仍未开始。
+31
View File
@@ -0,0 +1,31 @@
# #3 英语技术验证
独立小样,不是学习端正式功能;不连接 MySQL、不读取账号、不保存输入。仅本机、单进程串行运行,默认英语。长期契约与实测结论见 [开发验证 Wiki](https://git.ilapage.cn/OPC/lexgo/wiki/Local-Development-and-Verification)。
从仓库根目录运行(Windows PowerShell,需要 uv、Node):
```powershell
uv venv --python 3.12.12 .local/nlp-venv
uv pip install --python .local/nlp-venv/Scripts/python.exe -r spikes/english/requirements.lock
.local/nlp-venv/Scripts/python.exe spikes/english/setup_resources.py
uv pip install --python .local/nlp-venv/Scripts/python.exe --no-deps .local/nlp-resources/en_core_web_sm-3.8.0-py3-none-any.whl
.local/nlp-venv/Scripts/python.exe -m unittest discover -s spikes/english -v
node --test spikes/english/view.test.mjs
.local/nlp-venv/Scripts/python.exe spikes/english/benchmark.py
.local/nlp-venv/Scripts/python.exe spikes/english/app.py
```
打开 <http://127.0.0.1:5183/>。点击“分析文本”后点单词;`went` 应以 `go` 查询。手动查询仅查输入形式,不猜测词性:`dog` 直接命中,`went` 无结果。关闭进程即停止小样;端口占用时使用 `--port 5185`。用 `--resources .local/absent-resources` 启动可验证词典缺失,模型与词典独立加载。
准备依赖和资源时需要联网;安装完成后运行不依赖在线翻译或下载服务。`test_engine.py` 与 `benchmark.py` 禁止 socket connect,用真实模型与词典验证离线运行。网络下载失败可重新执行准备命令,已有资源先校验再复用。默认系统 Python 3.8 不适用,命令必须使用上述独立环境。
资源版本、固定下载 URL 和 SHA256 见 `resources.json`,Python 依赖固定于 `requirements.lock`。大文件只保存在忽略的 `.local/nlp-resources`。
许可与来源:
- [spaCy 3.8.7](https://pypi.org/pypi/spacy/3.8.7/json) 与 [en_core_web_sm 3.8.0](https://github.com/explosion/spacy-models/releases/tag/en_core_web_sm-3.8.0):MIT;安装包保留其许可证。模型的 POS/lemma 组件保留,parser/NER 在本小样中停用;不输出句界。
- [NLTK](https://github.com/nltk/nltk/blob/3.9.2/LICENSE.txt):Apache-2.0;只读取本地词典文件,无隐式 downloader。
- [Princeton WordNet 3.0](https://wordnet.princeton.edu/license-and-commercial-use):WordNet 3.0 许可证;下载 ZIP 完整保留 `wordnet/LICENSE`、版权及免责声明。英英释义,不提供中文翻译。再分发必须保留许可声明。
- [WordNet 原生格式](https://wordnet.princeton.edu/documentation/wndb5wn) 是 index/data/exception 文件,本小样直接读取 ZIP 中原始文件,不使用 go-admin 的系统枚举字典。
`benchmark-result.json` 是本机虚构语料的测量样本,不代表一般准确率或生产性能承诺。
+69
View File
@@ -0,0 +1,69 @@
import {validateTokens,latestOnly} from './view.mjs'
const source=document.querySelector('#source'), reading=document.querySelector('#reading'), definition=document.querySelector('#definition'), notice=document.querySelector('#notice'), position=document.querySelector('#position'), query=document.querySelector('#query'), analyzeButton=document.querySelector('#analyze')
const analysis=latestOnly(), lookup=latestOnly()
const labels={exact:'直接命中',lemma:'按原形查询',not_found:'未找到释义',resource_missing:'本地资源缺失',invalid_input:'输入格式无效或文本超过限制',internal_error:'暂时无法处理,请重试',too_large:'文本超过限制'}
async function post(path,payload) {
const response=await fetch(path,{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify(payload)})
const data=await response.json()
if(!response.ok) throw new Error(labels[data.status] || '请求失败,请重试')
return data
}
function element(tag,text) {const node=document.createElement(tag);node.textContent=text;return node}
async function showWord(surface,lemma='',token=null) {
const version=lookup.next()
query.value=surface
definition.replaceChildren(element('p','查询中…'))
position.textContent=token ? JSON.stringify({text:token.text,lemma:token.lemma,code_point:[token.start_cp,token.end_cp],utf8:[token.start_utf8,token.end_utf8],utf16:[token.start_utf16,token.end_utf16]},null,2) : '手动查询,无原文位置'
try {
const data=await post('/lookup',{surface,lemma})
if(!lookup.current(version)) return
definition.replaceChildren(element('h3',data.matched_form || surface),element('small',labels[data.status] || data.status))
if(data.entries.length) {
const list=document.createElement('ol')
for(const entry of data.entries) {
const item=element('li',entry.definition)
item.prepend(element('small',entry.pos+' · '))
for(const example of entry.examples) item.append(element('p',example))
list.append(item)
}
definition.append(list)
}
} catch(error) {if(lookup.current(version)) definition.textContent=error.message}
}
function resetReading() {
lookup.next()
reading.replaceChildren()
definition.textContent='点击文中的单词。'
position.textContent='尚未选择单词'
}
async function analyze() {
const version=analysis.next(), text=source.value
resetReading()
analyzeButton.disabled=true
notice.textContent='正在分析…'
try {
const result=await post('/analyze',{text})
if(!analysis.current(version)) return
validateTokens(text,result.tokens)
const fragment=document.createDocumentFragment()
for(const token of result.tokens) {
if(token.kind==='word') {
const button=element('button',token.text)
button.type='button'
button.addEventListener('click',()=>{
reading.querySelector('.selected')?.classList.remove('selected')
button.classList.add('selected')
showWord(token.text,token.lemma,token)
})
fragment.append(button)
} else fragment.append(document.createTextNode(token.text))
}
reading.replaceChildren(fragment)
notice.textContent=text ? '已分析 · 点击单词查词' : '请输入阅读文本'
} catch(error) {if(analysis.current(version)) notice.textContent=error.message}
finally {if(analysis.current(version)) analyzeButton.disabled=false}
}
source.addEventListener('input',()=>{analysis.next();resetReading();analyzeButton.disabled=false;notice.textContent='文本已修改,请重新分析'})
analyzeButton.addEventListener('click',analyze)
document.querySelector('#lookup-form').addEventListener('submit',event=>{event.preventDefault();showWord(query.value)})
analyze()
+117
View File
@@ -0,0 +1,117 @@
"""Loopback-only, ephemeral English experiment. Not a production API."""
import argparse
import json
from http.server import BaseHTTPRequestHandler, HTTPServer
from pathlib import Path
STATIC = Path(__file__).parent
ROOT = STATIC.parent.parent
def create_server(engine, port=5183):
class Handler(BaseHTTPRequestHandler):
def setup(self):
super().setup()
self.connection.settimeout(10)
def log_message(self, *_args):
pass # Do not retain input or query text in access logs.
def reply(self, code, body, mime='application/json; charset=utf-8'):
if not isinstance(body, bytes):
body = json.dumps(body, ensure_ascii=True).encode('utf-8')
self.send_response(code)
self.send_header('Content-Type', mime)
self.send_header('Content-Length', str(len(body)))
self.send_header('Cache-Control', 'no-store')
self.send_header('X-Content-Type-Options', 'nosniff')
self.send_header('Content-Security-Policy', "default-src 'self'; style-src 'self'; script-src 'self'; connect-src 'self'; frame-ancestors 'none'; base-uri 'none'")
self.end_headers()
self.wfile.write(body)
def local_request(self):
host = self.headers.get('Host')
allowed = {f'127.0.0.1:{self.server.server_port}', f'localhost:{self.server.server_port}'}
origin = self.headers.get('Origin')
if host not in allowed or (origin is not None and origin not in {'http://' + h for h in allowed}):
self.reply(403, {'status': 'forbidden'})
return False
return True
def do_GET(self):
if not self.local_request():
return
files = {'/': ('index.html', 'text/html'), '/app.mjs': ('app.mjs', 'text/javascript'), '/view.mjs': ('view.mjs', 'text/javascript'), '/style.css': ('style.css', 'text/css')}
if self.path not in files:
self.reply(404, {'status': 'not_found'})
return
name, mime = files[self.path]
self.reply(200, (STATIC / name).read_bytes(), mime + '; charset=utf-8')
def do_POST(self):
if not self.local_request():
return
if self.path not in {'/analyze', '/lookup'}:
self.reply(404, {'status': 'not_found'})
return
try:
length = int(self.headers.get('Content-Length', '0'))
except ValueError:
self.reply(400, {'status': 'invalid_input'})
return
if length > 1_000_000:
self.reply(413, {'status': 'too_large'})
return
if length <= 0:
self.reply(400, {'status': 'invalid_input'})
return
if self.headers.get_content_type() != 'application/json':
self.reply(415, {'status': 'invalid_content_type'})
return
try:
payload = json.loads(self.rfile.read(length))
if not isinstance(payload, dict):
raise ValueError('object required')
key = 'text' if self.path == '/analyze' else 'surface'
if not isinstance(payload.get(key), str):
raise ValueError('string required')
if self.path == '/analyze':
result = engine.analyze(payload['text'])
else:
lemma = payload.get('lemma', '')
if not isinstance(lemma, str) or len(lemma) > 200 or len(payload['surface']) > 200:
raise ValueError('invalid query')
result = engine.lookup(payload['surface'], lemma)
self.reply(200, result)
except (ValueError, TypeError, UnicodeError):
self.reply(400, {'status': 'invalid_input'})
except Exception as exc:
# ResourceMissing is deliberately exposed without filesystem paths.
from engine import ResourceMissing
if isinstance(exc, ResourceMissing):
self.reply(503, {'status': 'resource_missing', 'resource': exc.resource})
else:
self.reply(500, {'status': 'internal_error'})
server = HTTPServer(('127.0.0.1', port), Handler)
return server
def main():
from engine import Engine
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('--port', type=int, default=5183)
parser.add_argument('--resources', type=Path, default=ROOT / '.local/nlp-resources')
args = parser.parse_args()
server = create_server(Engine(args.resources), args.port)
print(f'LexGo English experiment: http://127.0.0.1:{server.server_port}/', flush=True)
try:
server.serve_forever()
except KeyboardInterrupt:
pass
finally:
server.server_close()
if __name__ == '__main__':
main()
+181
View File
@@ -0,0 +1,181 @@
{
"measured_at_utc": "2026-09-10T12:55:12.016850+00:00",
"environment": {
"python": "3.12.12",
"platform": "Windows-10-10.0.19044-SP0",
"processor": "Intel64 Family 6 Model 140 Stepping 1, GenuineIntel",
"logical_cpus": 8,
"packages": {
"spacy": "3.8.7",
"nltk": "3.9.2",
"en-core-web-sm": "3.8.0"
}
},
"offline": true,
"cold_engine_load_ms": 2355.1940000616014,
"text_codepoints": 100000,
"text_utf8_bytes": 106095,
"quality": {
"cases": [
{
"sentence": "She went home.",
"surface": "went",
"expected": "go",
"spacy": "go",
"baseline": "go",
"spacy_correct": true,
"baseline_correct": true
},
{
"sentence": "The children ate apples.",
"surface": "children",
"expected": "child",
"spacy": "child",
"baseline": "child",
"spacy_correct": true,
"baseline_correct": true
},
{
"sentence": "The children ate apples.",
"surface": "ate",
"expected": "eat",
"spacy": "eat",
"baseline": "ate",
"spacy_correct": true,
"baseline_correct": false
},
{
"sentence": "The dogs ran quickly.",
"surface": "dogs",
"expected": "dog",
"spacy": "dog",
"baseline": "dog",
"spacy_correct": true,
"baseline_correct": true
},
{
"sentence": "The dogs ran quickly.",
"surface": "ran",
"expected": "run",
"spacy": "run",
"baseline": "run",
"spacy_correct": true,
"baseline_correct": true
},
{
"sentence": "I saw a bird.",
"surface": "saw",
"expected": "see",
"spacy": "see",
"baseline": "saw",
"spacy_correct": true,
"baseline_correct": false
},
{
"sentence": "The saw is sharp.",
"surface": "saw",
"expected": "saw",
"spacy": "saw",
"baseline": "saw",
"spacy_correct": true,
"baseline_correct": true
},
{
"sentence": "She leaves today.",
"surface": "leaves",
"expected": "leave",
"spacy": "leave",
"baseline": "leaf",
"spacy_correct": true,
"baseline_correct": false
},
{
"sentence": "The leaves fell.",
"surface": "leaves",
"expected": "leaf",
"spacy": "leave",
"baseline": "leaf",
"spacy_correct": false,
"baseline_correct": true
},
{
"sentence": "They are reading books.",
"surface": "reading",
"expected": "read",
"spacy": "read",
"baseline": "reading",
"spacy_correct": true,
"baseline_correct": false
},
{
"sentence": "He was better yesterday.",
"surface": "was",
"expected": "be",
"spacy": "be",
"baseline": "wa",
"spacy_correct": true,
"baseline_correct": false
},
{
"sentence": "She can't go.",
"surface": "n't",
"expected": "not",
"spacy": "not",
"baseline": null,
"spacy_correct": true,
"baseline_correct": false
}
],
"total": 12,
"spacy_correct": 11,
"baseline_correct": 6
},
"measurements": {
"first_analysis": {
"repeats": 1,
"median_ms": 2.9011000879108906,
"p95_ms": 2.9011000879108906
},
"first_lookup": {
"repeats": 1,
"median_ms": 78.619199921377,
"p95_ms": 78.619199921377
},
"spacy_100k": {
"repeats": 3,
"median_ms": 1432.92090005707,
"p95_ms": 1499.0439999382943,
"codepoints_per_second": 69787.52281163407
},
"baseline_100k": {
"repeats": 3,
"median_ms": 336.1744999419898,
"p95_ms": 343.62249996047467,
"codepoints_per_second": 297464.56086721626
},
"lookup": {
"dog": {
"repeats": 100,
"median_ms": 0.014899997040629387,
"p95_ms": 0.022999942302703857
},
"went": {
"repeats": 100,
"median_ms": 0.021250045392662287,
"p95_ms": 0.02929999027401209
},
"zzzxqvfiction": {
"repeats": 100,
"median_ms": 0.003600027412176132,
"p95_ms": 0.004999921657145023
}
}
},
"limitations": [
"Synthetic microbenchmark; 12 selected cases do not establish general accuracy.",
"Baseline excludes UTF-8/UTF-16 conversion; spaCy timing includes full analyze contract.",
"Repeated queries are warm-process; no claim about production concurrency.",
"Cold engine load includes dependency imports but OS filesystem caches may be warm.",
"POS and sense disambiguation are not provided by dictionary lookup."
]
}
+112
View File
@@ -0,0 +1,112 @@
"""Offline synthetic measurements; stdout is one JSON document, no text input log."""
import argparse
from datetime import datetime, timezone
import importlib.metadata
import json
import math
import os
from pathlib import Path
import platform
import re
import statistics
import time
from unittest.mock import patch
from engine import Engine
CASES = [
('She went home.', 'went', 'go'),
('The children ate apples.', 'children', 'child'),
('The children ate apples.', 'ate', 'eat'),
('The dogs ran quickly.', 'dogs', 'dog'),
('The dogs ran quickly.', 'ran', 'run'),
('I saw a bird.', 'saw', 'see'),
('The saw is sharp.', 'saw', 'saw'),
('She leaves today.', 'leaves', 'leave'),
('The leaves fell.', 'leaves', 'leaf'),
('They are reading books.', 'reading', 'read'),
('He was better yesterday.', 'was', 'be'),
("She can't go.", "n't", 'not'),
]
def summary(samples):
values = sorted(samples)
return {'repeats': len(values), 'median_ms': statistics.median(values) * 1000,
'p95_ms': values[max(0, math.ceil(len(values) * .95) - 1)] * 1000}
def measure(function, repeats):
samples = []
for _ in range(repeats):
start = time.perf_counter()
function()
samples.append(time.perf_counter() - start)
return summary(samples)
def run(args):
start = time.perf_counter()
engine = Engine(args.resources)
load_seconds = time.perf_counter() - start
if engine.nlp is None or engine.wordnet is None:
raise RuntimeError(f'Resources unavailable: model={engine.model_error}, wordnet={engine.wordnet_error}')
measured = {}
measured['first_analysis'] = measure(lambda: engine.analyze('She went home.'), 1)
measured['first_lookup'] = measure(lambda: engine.lookup('dog'), 1)
baseline_pattern = re.compile(r"\w+(?:['’]\w+)*|\s+|[^\w\s]", re.UNICODE)
# Lower-cost baseline: regex spans and context-free WordNet morphology.
def baseline(text):
return [(match.group(), engine.wordnet.morphy(match.group().lower()) or match.group().lower(),
match.start(), match.end()) for match in baseline_pattern.finditer(text)]
quality = []
for sentence, surface, expected in CASES:
contextual = next((token['lemma'] for token in engine.analyze(sentence)['tokens']
if token['text'] == surface), None)
simple = next((lemma for token, lemma, _, _ in baseline(sentence) if token == surface), None)
quality.append(dict(sentence=sentence, surface=surface, expected=expected,
spacy=contextual, baseline=simple,
spacy_correct=contextual == expected, baseline_correct=simple == expected))
seed = "She went home. The children ate apples. I saw a bird. The leaves fell. Café 😀 e\u0301\r\n"
text = (seed * (100000 // len(seed) + 1))[:100000]
for name, function in [('spacy', engine.analyze), ('baseline', baseline)]:
metrics = measure(lambda: function(text), args.text_repeats)
metrics['codepoints_per_second'] = len(text) / (metrics['median_ms'] / 1000)
measured[name + '_100k'] = metrics
queries = [('dog', ''), ('went', 'go'), ('zzzxqvfiction', '')]
measured['lookup'] = {surface: measure(lambda: engine.lookup(surface, lemma), args.query_repeats)
for surface, lemma in queries}
return {
'measured_at_utc': datetime.now(timezone.utc).isoformat(),
'environment': {'python': platform.python_version(), 'platform': platform.platform(),
'processor': platform.processor(), 'logical_cpus': os.cpu_count(),
'packages': {name: importlib.metadata.version(name)
for name in ('spacy', 'nltk', 'en-core-web-sm')}},
'offline': True, 'cold_engine_load_ms': load_seconds * 1000,
'text_codepoints': len(text), 'text_utf8_bytes': len(text.encode()),
'quality': {'cases': quality, 'total': len(quality),
'spacy_correct': sum(case['spacy_correct'] for case in quality),
'baseline_correct': sum(case['baseline_correct'] for case in quality)},
'measurements': measured,
'limitations': [
'Synthetic microbenchmark; 12 selected cases do not establish general accuracy.',
'Baseline excludes UTF-8/UTF-16 conversion; spaCy timing includes full analyze contract.',
'Repeated queries are warm-process; no claim about production concurrency.',
'Cold engine load includes dependency imports but OS filesystem caches may be warm.',
'POS and sense disambiguation are not provided by dictionary lookup.',
],
}
if __name__ == '__main__':
parser = argparse.ArgumentParser()
parser.add_argument('--resources', type=Path, default=Path(__file__).resolve().parents[2] / '.local/nlp-resources')
parser.add_argument('--text-repeats', type=int, default=3)
parser.add_argument('--query-repeats', type=int, default=100)
args = parser.parse_args()
if args.text_repeats < 1 or args.query_repeats < 1:
parser.error('repeat counts must be positive')
with patch('socket.socket.connect', side_effect=AssertionError('network forbidden')):
print(json.dumps(run(args), ensure_ascii=True, indent=2))
+116
View File
@@ -0,0 +1,116 @@
"""Isolated English experiment. Resources are local; no downloader is used."""
import hashlib
from pathlib import Path
import unicodedata
import warnings
import zipfile
class ResourceMissing(RuntimeError):
def __init__(self, resource):
self.resource = resource
super().__init__(f'Local {resource} resource is unavailable')
def validate_text(text):
if not isinstance(text, str):
raise TypeError('text must be a string')
if len(text) > 100000:
raise ValueError('text exceeds 100000 code points')
if any(0xD800 <= ord(char) <= 0xDFFF for char in text):
raise ValueError('text contains an unpaired surrogate')
def lookup_form(text):
validate_text(text)
return unicodedata.normalize('NFC', text.casefold()).replace('’', "'").replace('‘', "'")
class Engine:
def __init__(self, resource_dir: Path, model_name='en_core_web_sm'):
self.nlp = None
self.wordnet = None
self.model_error = None
self.wordnet_error = None
try:
import spacy
self.nlp = spacy.load(model_name, disable=['parser', 'ner'])
except (ImportError, OSError, ValueError) as error:
self.model_error = type(error).__name__
try:
from nltk.corpus.reader import WordNetCorpusReader
from nltk.data import ZipFilePathPointer
class EnglishWordNet30Reader(WordNetCorpusReader):
def map_wn(self, version='wordnet'):
# NLTK's default cross-version OMW mapping loads a global
# corpus. English-only WordNet 3.0 needs no such mapping.
if self.get_version() != '3.0':
raise ValueError('This experiment requires WordNet 3.0')
return None
root = ZipFilePathPointer(str(Path(resource_dir) / 'wordnet.zip'), 'wordnet/')
with warnings.catch_warnings():
warnings.filterwarnings('ignore', message='The multilingual functions are not available with this Wordnet version', category=UserWarning)
self.wordnet = EnglishWordNet30Reader(root, None)
except (ImportError, OSError, LookupError, ValueError, zipfile.BadZipFile) as error:
self.wordnet_error = type(error).__name__
def analyze(self, text):
validate_text(text)
if self.nlp is None:
raise ResourceMissing('model')
# Prefix tables make conversion linear even for long Unicode documents.
utf8 = [0]
utf16 = [0]
for char in text:
utf8.append(utf8[-1] + len(char.encode('utf-8')))
utf16.append(utf16[-1] + (2 if ord(char) > 0xFFFF else 1))
tokens = []
def append(start, end, lemma, kind):
tokens.append(dict(text=text[start:end], lemma=lemma, kind=kind,
start_cp=start, end_cp=end,
start_utf8=utf8[start], end_utf8=utf8[end],
start_utf16=utf16[start], end_utf16=utf16[end]))
cursor = 0
for token in self.nlp(text):
if token.idx > cursor:
append(cursor, token.idx, '', 'space')
end = token.idx + len(token.text)
kind = 'space' if token.is_space else 'punctuation' if token.is_punct else 'word'
append(token.idx, end, token.lemma_ if kind == 'word' else '', kind)
cursor = end
if cursor < len(text):
append(cursor, len(text), '', 'space')
return dict(status='ok', contract_version='english-spike-v1', original_text=text,
text_sha256=hashlib.sha256(text.encode('utf-8')).hexdigest(), tokens=tokens)
def _exact_entries(self, form):
# NLTK 3.9.2 synsets() applies morphy even with check_exceptions=False.
# Read its loaded index directly to keep exact and explicit lemma distinct.
index = self.wordnet._lemma_pos_offset_map.get(form, {})
entries = []
for pos in ('n', 'v', 'a', 'r'):
for offset in index.get(pos, []):
synset = self.wordnet.synset_from_pos_and_offset(pos, offset)
entries.append(dict(lemma=form, pos=synset.pos(),
definition=synset.definition(), examples=synset.examples()))
if len(entries) == 12:
return entries
return entries
def lookup(self, surface, lemma=''):
form = lookup_form(surface)
fallback = lookup_form(lemma)
result = dict(query=surface, matched_form=None, entries=[])
if self.wordnet is None:
return dict(result, status='resource_missing', resource='wordnet')
for candidate, status in ((form, 'exact'), (fallback, 'lemma')):
if not candidate:
continue
entries = self._exact_entries(candidate)
if entries:
return dict(result, status=status, matched_form=candidate, entries=entries)
return dict(result, status='not_found')
+12
View File
@@ -0,0 +1,12 @@
<!doctype html>
<html lang="zh-CN">
<head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>LexGo · 英语验证</title><link rel="stylesheet" href="/style.css"></head>
<body>
<header><a class="brand" href="/">LexGo<span>.</span></a><span class="badge">英语 · 技术验证</span></header>
<main>
<section class="input-section"><label for="source">阅读文本</label><textarea id="source" spellcheck="false" maxlength="100000">Mira’s well-known dogs went home. She can't wait.
The children were running beside a café. 🙂 Café!</textarea><div class="actions"><span>仅本机处理 · 不保存</span><button id="analyze">分析文本</button></div></section>
<p id="notice" role="status" aria-live="polite"></p>
<div class="workspace"><section class="paper" aria-label="阅读结果"><h1>阅读</h1><div id="reading"></div></section><aside aria-label="词典"><div class="dictionary-header"><h2>本地词典</h2><span>WordNet 3.0</span></div><form id="lookup-form"><label class="sr-only" for="query">查询单词</label><input id="query" maxlength="200" placeholder="输入英语单词" autocomplete="off"><button>查询</button></form><div id="definition" aria-live="polite">点击文中的单词。</div><details><summary>原文位置</summary><pre id="position">尚未选择单词</pre></details></aside></div>
</main><footer>独立验证小样 · 英语释义</footer><script type="module" src="/app.mjs"></script>
</body></html>
+47
View File
@@ -0,0 +1,47 @@
annotated-doc==0.0.5
annotated-types==0.8.0
blis==1.3.3
catalogue==2.0.10
certifi==2026.7.22
charset-normalizer==3.5.1
click==8.5.0
cloudpathlib==0.25.0
cloudpickle==3.1.2
colorama==0.4.6
confection==0.1.5
cymem==2.0.13
idna==3.19
jinja2==3.1.6
joblib==1.6.0
langcodes==3.5.1
markdown-it-py==4.2.0
markupsafe==3.0.3
mdurl==0.1.2
murmurhash==1.0.15
nltk==3.9.2
numpy==2.5.3
packaging==26.3
preshed==3.0.13
pydantic==2.13.5
pydantic-core==2.46.5
pygments==2.21.0
regex==2026.9.10
requests==2.34.2
rich==15.0.0
setuptools==84.0.0
shellingham==1.5.4
smart-open==7.7.1
spacy==3.8.7
spacy-legacy==3.0.12
spacy-loggers==1.0.5
srsly==2.5.3
thinc==8.3.11
tqdm==4.70.0
typer==0.27.2
typer-slim==0.24.0
typing-extensions==4.16.0
typing-inspection==0.4.4
urllib3==2.7.0
wasabi==1.1.3
weasel==0.4.3
wrapt==2.4.0
+21
View File
@@ -0,0 +1,21 @@
{
"python": "3.12.12",
"spacy": "3.8.7",
"nltk": "3.9.2",
"model": {
"name": "en_core_web_sm",
"version": "3.8.0",
"file": "en_core_web_sm-3.8.0-py3-none-any.whl",
"url": "https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl",
"sha256": "1932429db727d4bff3deed6b34cfc05df17794f4a52eeb26cf8928f7c1a0fb85",
"license": "MIT"
},
"dictionary": {
"name": "Princeton WordNet",
"version": "3.0",
"file": "wordnet.zip",
"url": "https://raw.githubusercontent.com/nltk/nltk_data/96f9b3252457a2b97e52aec64c3dfceeb5c312d5/packages/corpora/wordnet.zip",
"sha256": "cbda5ea6eef7f36a97a43d4a75f85e07fccbb4f23657d27b4ccbc93e2646ab59",
"license": "WordNet 3.0 License (included in ZIP: wordnet/LICENSE)"
}
}
+30
View File
@@ -0,0 +1,30 @@
"""Explicit network-only preparation; runtime never calls this script."""
import hashlib
import json
from pathlib import Path
import urllib.request
HERE = Path(__file__).resolve().parent
DEST = HERE.parent.parent / '.local/nlp-resources'
def main():
manifest = json.loads((HERE / 'resources.json').read_text(encoding='utf-8'))
DEST.mkdir(parents=True, exist_ok=True)
for key in ('model', 'dictionary'):
item = manifest[key]
target = DEST / item['file']
if target.exists() and hashlib.sha256(target.read_bytes()).hexdigest() == item['sha256']:
print(key + ': checksum verified', flush=True)
continue
req = urllib.request.Request(item['url'], headers={'User-Agent': 'LexGo-English-Spike/1'})
with urllib.request.urlopen(req, timeout=120) as response:
data = response.read(32 * 1024 * 1024 + 1)
if hashlib.sha256(data).hexdigest() != item['sha256']:
raise RuntimeError(key + ': checksum mismatch; resource not installed')
target.write_bytes(data)
print(key + ': downloaded and checksum verified', flush=True)
if __name__ == '__main__':
main()
+1
View File
@@ -0,0 +1 @@
:root{color:#253b32;background:#f5f4ef;font-family:"Segoe UI","Microsoft YaHei",sans-serif;font-synthesis:none}*{box-sizing:border-box}body{margin:0}header{height:78px;border-bottom:1px solid #dfe3da;padding:0 5%;display:flex;align-items:center;justify-content:space-between;background:#fff}.brand{font-size:29px;font-weight:700;text-decoration:none;color:inherit;letter-spacing:-1px}.brand span{color:#3b8060}.badge{font-size:13px;color:#66776a}main{max-width:1260px;margin:35px auto;padding:0 28px}.input-section{border-bottom:1px solid #d5ddd4;padding-bottom:25px}label{display:block;font-weight:600;margin-bottom:12px}textarea{display:block;width:100%;min-height:135px;resize:vertical;border:1px solid #cbd5cb;border-radius:8px;background:#fff;padding:16px;color:#263b30;font:18px/1.7 Georgia,serif}textarea:focus,input:focus,button:focus-visible{outline:2px solid #538967;outline-offset:3px}.actions{display:flex;justify-content:space-between;align-items:center;margin-top:12px}.actions span,footer{font-size:12px;color:#738075}button{background:#2c6347;color:white;border:0;border-radius:5px;padding:10px 20px;cursor:pointer;font:inherit}button:disabled{opacity:.55;cursor:wait}#notice{font-size:14px;min-height:20px}.workspace{display:grid;grid-template-columns:minmax(0,1fr) 330px;gap:25px}.paper,aside{background:#fff;border:1px solid #e0e5dc;border-radius:8px}.paper{padding:28px 32px;min-height:340px}h1{font-size:13px;letter-spacing:2px;color:#6f7e72;margin:0 0 26px}#reading{font:23px/1.95 Georgia,"Times New Roman",serif;white-space:pre-wrap;overflow-wrap:anywhere}#reading button{font:inherit;color:inherit;padding:0;border-radius:2px;background:transparent;text-align:left}#reading button:hover,#reading button.selected{background:#e1edcf;box-shadow:0 2px #658447}aside{padding:24px}.dictionary-header{display:flex;justify-content:space-between;align-items:center;margin-bottom:20px}.dictionary-header h2{font-size:17px;margin:0}.dictionary-header span{font-size:11px;color:#7e887f}form{display:flex;gap:7px;margin-bottom:22px}input{min-width:0;width:100%;padding:9px;border:1px solid #cbd5cb;border-radius:4px;font:inherit}form button{padding:9px 12px;white-space:nowrap}#definition{font-size:14px;line-height:1.7;overflow-wrap:anywhere}#definition h3{font:27px Georgia,serif;margin:0 0 8px}#definition ol{padding-left:21px}#definition li{margin-bottom:13px}#definition small{color:#637567}details{margin-top:25px;border-top:1px solid #e1e6de;padding-top:15px;color:#7a847d;font-size:12px}summary{cursor:pointer}pre{white-space:pre-wrap;overflow-wrap:anywhere;font-size:11px}footer{text-align:center;padding:35px}.sr-only{position:absolute;width:1px;height:1px;overflow:hidden;clip-path:inset(50%)}@media(max-width:750px){main{padding:0 16px;margin-top:20px}.workspace{grid-template-columns:1fr}.paper{padding:24px;min-height:230px}#reading{font-size:21px}.actions span{font-size:11px}}
+72
View File
@@ -0,0 +1,72 @@
import http.client
import json
import threading
import unittest
from unittest.mock import Mock
from app import create_server
class HTTPTests(unittest.TestCase):
def setUp(self):
self.engine = Mock()
self.engine.analyze.return_value = {'status': 'ok', 'tokens': []}
self.engine.lookup.return_value = {'status': 'not_found', 'entries': []}
self.server = create_server(self.engine, 0)
self.thread = threading.Thread(target=self.server.serve_forever, daemon=True)
self.thread.start()
self.port = self.server.server_port
def tearDown(self):
self.server.shutdown()
self.server.server_close()
self.thread.join()
def call(self, method, path, body=None, headers=None):
c = http.client.HTTPConnection('127.0.0.1', self.port, timeout=3)
c.request(method, path, body, headers or {})
r = c.getresponse()
result = r.status, r.read(), dict(r.getheaders())
c.close()
return result
def test_local_page_and_no_arbitrary_file_access(self):
status, body, headers = self.call('GET', '/')
self.assertEqual(status, 200)
self.assertIn(b'LexGo', body)
self.assertIn('Content-Security-Policy', headers)
self.assertEqual(self.call('GET', '/../../.env.local')[0], 404)
def test_json_analyze_and_lookup(self):
self.assertEqual(self.call('POST', '/analyze', json.dumps({'text': 'Hello'}), {'Content-Type': 'application/json'})[0], 200)
self.engine.analyze.assert_called_once_with('Hello')
self.assertEqual(self.call('POST', '/lookup', json.dumps({'surface': 'went', 'lemma': 'go'}), {'Content-Type': 'application/json'})[0], 200)
self.engine.lookup.assert_called_once_with('went', 'go')
def test_reject_cross_origin_and_rebinding(self):
for headers in ({'Origin': 'https://evil.example'}, {'Host': 'evil.example'}):
self.assertEqual(self.call('POST', '/analyze', '{}', headers)[0], 403)
self.engine.analyze.assert_not_called()
def test_invalid_payload_and_size(self):
for payload in ('[]', '{}', '{', '{"text":42}'):
self.assertEqual(self.call('POST', '/analyze', payload, {'Content-Type': 'application/json'})[0], 400)
self.assertEqual(self.call('POST', '/analyze', '{}', {'Content-Type': 'text/plain'})[0], 415)
self.assertEqual(self.call('POST', '/analyze', '{}', {'Content-Type': 'application/json', 'Content-Length': '1000001'})[0], 413)
def test_internal_errors_do_not_echo_input(self):
self.engine.analyze.side_effect = RuntimeError('private sample')
status, body, _ = self.call('POST', '/analyze', '{"text":"x"}', {'Content-Type': 'application/json'})
self.assertEqual(status, 500)
self.assertNotIn(b'private sample', body)
def test_missing_model_is_distinct_from_invalid_input(self):
from engine import ResourceMissing
self.engine.analyze.side_effect = ResourceMissing('model')
status, body, _ = self.call('POST', '/analyze', '{"text":"x"}', {'Content-Type': 'application/json'})
self.assertEqual(status, 503)
self.assertEqual(json.loads(body), {'status': 'resource_missing', 'resource': 'model'})
if __name__ == '__main__':
unittest.main()
+92
View File
@@ -0,0 +1,92 @@
"""Run with the isolated Python: -m unittest discover -s spikes/english -v."""
import hashlib
from pathlib import Path
import tempfile
import unittest
from unittest.mock import patch
try:
from engine import Engine, ResourceMissing
except ImportError:
Engine = None
RESOURCES = Path(__file__).resolve().parents[2] / '.local/nlp-resources'
class EngineTests(unittest.TestCase):
@classmethod
def setUpClass(cls):
cls.network = patch('socket.socket.connect', side_effect=AssertionError('network forbidden'))
cls.network.start()
cls.addClassCleanup(cls.network.stop)
if Engine:
cls.engine = Engine(RESOURCES)
def setUp(self):
self.assertIsNotNone(Engine, 'English engine has not been implemented')
def test_unicode_partition_and_three_offsets(self):
for text in ['', ' \t\r\n', " She went!\r\nDogs’ paws\tcan't. e\u0301 café 😀 中文\u00a0\u200bend ",
'well-known mother-in-law 👩‍💻 👨‍👩‍👧‍👦 🏳️‍🌈']:
with self.subTest(text=text):
result = self.engine.analyze(text)
self.assertEqual(result['status'], 'ok')
self.assertEqual(result['contract_version'], 'english-spike-v1')
self.assertEqual(result['original_text'], text)
self.assertEqual(result['text_sha256'], hashlib.sha256(text.encode()).hexdigest())
tokens = result['tokens']
self.assertEqual(''.join(t['text'] for t in tokens), text)
cursor = 0
for token in tokens:
self.assertEqual(token['start_cp'], cursor)
cursor = token['end_cp']
self.assertGreater(cursor, token['start_cp'])
self.assertEqual(text[token['start_cp']:cursor], token['text'])
for encoding, unit, suffix in [('utf-8', 1, 'utf8'), ('utf-16-le', 2, 'utf16')]:
start, end = token['start_' + suffix], token['end_' + suffix]
self.assertEqual(text.encode(encoding)[start*unit:end*unit].decode(encoding), token['text'])
self.assertEqual(len(text[:token['start_cp']].encode(encoding)) // unit, start)
self.assertIn(token['kind'], ['word', 'space', 'punctuation'])
self.assertEqual(cursor, len(text))
def test_contextual_irregular_lemma(self):
tokens = self.engine.analyze('She went home. The children ate apples.')['tokens']
lemmas = {t['text']: t['lemma'] for t in tokens}
self.assertEqual(lemmas['went'], 'go')
self.assertEqual(lemmas['children'], 'child')
self.assertEqual(lemmas['ate'], 'eat')
def test_exact_then_explicit_lemma(self):
exact = self.engine.lookup('DOG')
self.assertEqual(exact['status'], 'exact')
self.assertEqual(exact['matched_form'], 'dog')
self.assertTrue(exact['entries'])
self.assertLessEqual(len(exact['entries']), 12)
self.assertEqual(exact, self.engine.lookup('DOG'))
self.assertEqual(self.engine.lookup('went')['status'], 'not_found')
lemma = self.engine.lookup('went', 'go')
self.assertEqual(lemma['status'], 'lemma')
self.assertEqual(lemma['matched_form'], 'go')
self.assertEqual(self.engine.lookup('zzzxqvfiction')['status'], 'not_found')
def test_validation(self):
for text in ['x' * 100001, '\ud800']:
with self.assertRaises(ValueError):
self.engine.analyze(text)
with self.assertRaises(TypeError):
self.engine.analyze(None)
with self.assertRaises(ValueError):
self.engine.lookup('\udfff')
def test_missing_resources_are_not_misses(self):
with tempfile.TemporaryDirectory() as directory:
engine = Engine(Path(directory), model_name='nonexistent_english_spike_model')
self.assertEqual(engine.lookup('dog')['status'], 'resource_missing')
self.assertEqual(engine.lookup('dog')['resource'], 'wordnet')
with self.assertRaises(ResourceMissing) as error:
engine.analyze('dog')
self.assertEqual(error.exception.resource, 'model')
if __name__ == '__main__':
unittest.main()
+14
View File
@@ -0,0 +1,14 @@
export function validateTokens(text, tokens) {
let end=0
for (const token of tokens) {
if (!Number.isInteger(token.start_utf16) || !Number.isInteger(token.end_utf16) || token.start_utf16!==end || token.end_utf16<=end || text.slice(token.start_utf16,token.end_utf16)!==token.text) throw new Error('原文位置校验失败')
end=token.end_utf16
}
if(end!==text.length) throw new Error('原文还原失败')
return true
}
export function latestOnly() {
let version=0
return {next:()=>++version,current:value=>value===version}
}
+19
View File
@@ -0,0 +1,19 @@
import test from 'node:test'
import assert from 'node:assert/strict'
import { validateTokens, latestOnly } from './view.mjs'
test('UTF-16 positions reconstruct emoji and combining characters without normalization', () => {
const text = '🙂 Café'
const tokens = [{text:'🙂',start_utf16:0,end_utf16:2},{text:' ',start_utf16:2,end_utf16:3},{text:'Café',start_utf16:3,end_utf16:8}]
assert.equal(validateTokens(text,tokens), true)
assert.throws(() => validateTokens(text,[{text:'🙂',start_utf16:0,end_utf16:1}]))
assert.throws(() => validateTokens(text,[]))
})
test('old analysis and lookup responses cannot replace newer text or selection', () => {
const gate=latestOnly()
const first=gate.next(), second=gate.next()
assert.equal(gate.current(first),false)
assert.equal(gate.current(second),true)
gate.next()
assert.equal(gate.current(second),false)
})