feat: 验证英语分词、原文定位与离线词典 (#3)
This commit is contained in:
@@ -258,7 +258,7 @@ MVP 内所有单元任务通过后才能做 MVP 集成验收;MVP 通过后才
|
||||
|
||||
- 治理模式:轻量。数据库:MySQL 8(用户于 2026-09-10 确认);具体小版本在工程验证后锁定。
|
||||
- 远端:https://git.ilapage.cn/OPC/lexgo.git;分支 main。不得把邻接 dev_harness 工作区当成本项目工作区。
|
||||
- 工程基础 #2 已通过用户验收:server 基于指定 go-admin 选用模型扩展账号/会话 API,admin 复用 go-admin-ui,learner 为独立 Vue 3 + TypeScript + Vite 工程。默认英语;阅读、导入、词典、复习及 Python NLP 尚未实现或验证。
|
||||
- 工程基础 #2 已通过用户验收:server 基于指定 go-admin 选用模型扩展账号/会话 API,admin 复用 go-admin-ui,learner 为独立 Vue 3 + TypeScript + Vite 工程。默认英语;阅读、导入、词典与复习尚未接入产品;#3 已完成独立 Python NLP/词典验证小样,待验收。
|
||||
- 原四份研究保留为历史参考;PostgreSQL 建议被 MySQL 8 决策覆盖,U/A/N 索引用于追踪而不是批准所有范围。
|
||||
- 用户/语言数据所有权、Unicode 原文位置、任务和复习幂等、完整备份恢复是后续方案的必要验收边界。
|
||||
- 当前 MCP 连接其他 Gitea 站点,需使用目标站点 API 时记录原因;凭据仅从安全配置进入进程。
|
||||
@@ -278,3 +278,4 @@ MVP 内所有单元任务通过后才能做 MVP 集成验收;MVP 通过后才
|
||||
- 已验证 MySQL 8.4.3,本机 127.0.0.1:3308;开发库 lexgo_dev、测试库 lexgo_test_issue2。密码只从环境或忽略的 .env.local 读取。迁移测试只能使用 lexgo_test_ 前缀专用库,不能借用其他数据库。
|
||||
- 后端命令使用 `python scripts/server.py migrate|bootstrap|serve|build|test|test-integration`;仅显式 migrate 修改表。bootstrap 只接受尚无账号的 LexGo 库,不覆盖已有管理员。Go 1.26.5、Node 22.22.1、pnpm 9.15.1;两端分别构建。
|
||||
- #18 登录日志与操作审计已通过用户验收:schema v2 显式迁移;日志只保存白名单字段,禁止保存凭据、请求/响应正文及私人学习内容。仅管理员查询,默认保留 90 天;启动/每小时及 `python scripts/server.py audit-cleanup` 仅清理两张审计表的过期记录。
|
||||
- #3 独立小样位于 `spikes/english/`,使用 `.local/nlp-venv/Scripts/python.exe`(3.12.12)运行;固定 spaCy 3.8.7、英语模型 3.8.0、NLTK 3.9.2、WordNet 3.0。资源仅显式准备时下载,摘要见 resources.json。不得把本机无账号的实验接口用于正式学习端;后续集成仍需 Go 授权、数据归属和任务设计。原文不归一化,位置区分 cp/UTF-8/UTF-16,lemma 不自动合并学习状态。
|
||||
|
||||
@@ -5,6 +5,7 @@
|
||||
已确认:**DevHarness 轻量模式、MySQL 8、go-admin 管理端**。工程基础 #2 已通过验收:两端用户名登录、学习账号管理、可撤销会话和本人英语空空间。管理端基于指定 go-admin/go-admin-ui 选用模块,学习端为独立 Vue 3 + TypeScript + Vite 工程,共用 Go 后端和 MySQL 8.4.3。#18 登录日志与操作审计已通过用户验收,支持管理员查询和 90 天保留清理。阅读、导入、词典与复习尚未实现。MVP 定位为“支持多账号、数据独立的自托管学习工具”,先邀请少量用户使用;F01~F12 已确认,X 系列后置。
|
||||
|
||||
- [文档入口](docs/README.md) · [线上 Wiki](https://git.ilapage.cn/OPC/lexgo/wiki/Home)
|
||||
- [英语分词与离线词典验证小样](spikes/english/README.md)(#3 待验收,独立本机入口)
|
||||
- [项目档案](docs/00-project-profile.md) · [需求总览](docs/09-product-requirements-overview.md)
|
||||
- [工作量估算](docs/10-workload-estimate.md):#2 验收后原范围剩余 46~75 人日,新增 #18 的 3~5 人日计划后为 49~80 人日;技术验证后重估,旧全量研究仅供参考。
|
||||
- [四阶段实施总览 #16](https://git.ilapage.cn/OPC/lexgo/issues/16):14 张单元工单,工程基础 → 技术验证 → 首条学习闭环 → 补齐 MVP;原型 v1 已获用户验收。两端使用账号(用户名)+密码登录,不要求邮箱。
|
||||
|
||||
@@ -2,8 +2,8 @@
|
||||
generated: true (请先修改 Gitea Wiki,禁止直接编辑本文件)
|
||||
wiki_page: Architecture-and-Code-Map
|
||||
wiki_url: https://git.ilapage.cn/OPC/lexgo/wiki/Architecture-and-Code-Map.-
|
||||
wiki_revision: 3b6f6db7f0fe1019b5811f70af8d6174cec0b950
|
||||
synchronized_at: 2026-09-10T12:36:57Z
|
||||
wiki_revision: 0806b7fa3ee8606020c9d7185e26c059c0b761c9
|
||||
synchronized_at: 2026-09-10T12:58:33Z
|
||||
<!-- gitea-wiki-mirror:end -->
|
||||
|
||||
# 架构与代码地图
|
||||
@@ -23,7 +23,7 @@ Go 承担业务与后台任务,浏览器提供阅读学习界面,NLP 保留
|
||||
| `dev_scripts/harness.py`、`dev_scripts/wiki_docs.py` | 原样复制的 DevHarness 工具 | 治理工具入口,无业务 API |
|
||||
| `tests/` | 上游治理工具与文档结构测试 | 不验证阅读、NLP 或 SRS |
|
||||
|
||||
不存在产品入口、数据库迁移或前端页面。拟定职责:identity(身份)、library(书库)、ingestion(导入)、lexicon(全局词典)、vocabulary(个人词语)、review(复习)、progress(统计)、administration(管理)。底座固定后再决定具体目录。
|
||||
工程入口与已实现模块见下方 #2/#18;学习领域拟定职责:identity(身份)、library(书库)、ingestion(导入)、lexicon(全局词典)、vocabulary(个人词语)、review(复习)、progress(统计)、administration(管理)。底座固定后再决定具体目录。
|
||||
|
||||
## 两条主要执行路径
|
||||
|
||||
@@ -47,7 +47,7 @@ Go 承担业务与后台任务,浏览器提供阅读学习界面,NLP 保留
|
||||
|
||||
## 学习端与管理端的目标架构
|
||||
|
||||
技术方向已纳入本轮方案,以下是目标结构,尚无对应产品目录或可运行应用。
|
||||
以下是目标结构;账号、管理端与学习空空间已由 #2 实现。英语 NLP 已完成 #3 独立验证,尚未接入业务 API。
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
@@ -56,7 +56,7 @@ flowchart TD
|
||||
API --> B[独立学习业务模块]
|
||||
B --> DB[(MySQL 8)]
|
||||
B --> W[后台任务:具体选型待验证]
|
||||
W --> N[NLP 服务:Python 方案待确认]
|
||||
W --> N[NLP:Python 小样已验证,生产集成待实施]
|
||||
```
|
||||
|
||||
| 交付部分 | 建设方式 | 复用与自建边界 |
|
||||
@@ -131,3 +131,20 @@ server/app/lexgo/audit.go 定义两类白名单字段日志、筛选分页、失
|
||||
database.go 显式迁移至 v2,两张新增表均以 created_at/id 建立排序清理索引,账号字段建查询索引,无业务表级联删除。cmd/lexgo/main.go 的服务进程在启动和每小时执行审计清理,每次最多运行一分钟、每批删除 1000 条,仅影响过期审计记录。
|
||||
|
||||
admin/src/views/AuditLogs.vue 通过 kind 复用登录/操作列表;audit-logs.mjs 负责筛选编码和请求序号,session.mjs 继续进行管理员及会话 generation 校验。切换页面/账号清空日志,普通翻页保留总数,防止分页组件跳回第一页。菜单与标题按当前路由显示。
|
||||
|
||||
|
||||
## 英语分词与本地词典验证(#3,待验收)
|
||||
|
||||
`spikes/english/` 是独立可运行验证小样,不是学习端生产功能。推荐后续采用 Python 3.12.12、spaCy 3.8.7、en_core_web_sm 3.8.0(保留 tok2vec/tagger/attribute_ruler/lemmatizer,停用 parser/ner)和 NLTK 3.9.2 读取 WordNet 3.0。Go 继续管理用户、权限、任务和持久数据,后续通过显式契约调用 NLP;本单未新增 Go API、MySQL 表或常驻部署实例。
|
||||
|
||||
| 文件 | 作用 |
|
||||
|---|---|
|
||||
| engine.py | 原文分词、lemma、三个位置单位、直接/lemma 查词;只读本地资源 |
|
||||
| app.py、index.html、app.mjs、view.mjs、style.css | loopback 临时 HTTP 小样、输入/阅读/查词结果;单进程串行,输入不落盘 |
|
||||
| resources.json、setup_resources.py、requirements.lock | 固定版本、来源和 SHA256;显式联网准备,运行期无自动下载 |
|
||||
| test_engine.py、test_app.py、view.test.mjs | 真实模型离线验证、HTTP 边界与浏览器偏移/迟到响应测试 |
|
||||
| benchmark.py、benchmark-result.json | 虚构语料的候选对照、长文/查询实测及环境样本 |
|
||||
|
||||
WordNet 使用 ZIP 内原始 index/data/exception 文件,不使用 SysDict 或新增业务库。NLTK 默认 synsets 会隐式词形还原,本小样直接读取其固定版本索引以区分 exact 和显式 lemma;禁用依赖全局 corpus 的 OMW 跨版本映射,只接受 WordNet 3.0。升级 NLTK 或词典时必须重跑契约测试。
|
||||
|
||||
候选比较:正则分词+WordNet 默认名词 morphology 依赖少、速度快,但不具备上下文判断,缩写和词性歧义处理弱;纯 Go 规则同样需要自行维护这些语言规则。本次 spaCy 在 12 个明确样例中答对 11 个,基线 6 个,因此推荐保留独立 Python NLP 边界。样例量不足以证明总体准确率;不宣称部署或正式阅读功能已完成。
|
||||
|
||||
@@ -2,8 +2,8 @@
|
||||
generated: true (请先修改 Gitea Wiki,禁止直接编辑本文件)
|
||||
wiki_page: Business-Rules-and-Glossary
|
||||
wiki_url: https://git.ilapage.cn/OPC/lexgo/wiki/Business-Rules-and-Glossary.-
|
||||
wiki_revision: 7061627fed9b5acd74fbd2369f27acadc0fd2e2a
|
||||
synchronized_at: 2026-09-10T12:36:59Z
|
||||
wiki_revision: 493544b5fd723c028d9173ecf590b66fea72c20a
|
||||
synchronized_at: 2026-09-10T12:58:34Z
|
||||
<!-- gitea-wiki-mirror:end -->
|
||||
|
||||
# 业务规则与术语
|
||||
@@ -80,3 +80,22 @@ M0 固定首发语言语料、词条身份规则、短语选择与重叠规则
|
||||
- 日志绝不保存密码、token、Cookie、请求/响应正文、错误堆栈或私人学习内容。合法账号、IP 属于本功能必要的审计数据,仅管理员可查询。
|
||||
- GET /api/v1/login-logs 与 /operation-logs:未登录 401,学习者 403;page 默认 1,limit 默认 20、最大 100。username 精确匹配;操作日志匹配操作人或目标账号。result 为 success/failure,action 限定枚举,from/to 为 RFC3339。返回 data.items/total/page/limit,按时间和编号倒序,时间按毫秒存储、浏览器按本地时区展示。
|
||||
- 固定保留最近 90 天,查询即排除过期记录;默认起止为保留边界和当前时间。清理只删除 created_at 严格早于边界的两表记录,不影响账号、空间、会话。无清空全部或导出按钮;本期不开放保留时长配置。
|
||||
|
||||
|
||||
## #3 英语位置与查询实验契约 v1
|
||||
|
||||
实验版本 `english-spike-v1`。POST /analyze 接收 {text},返回 status、contract_version、original_text、text_sha256 和 tokens;仅为本机小样接口,不是生产 API。原文以收到的字符串为准,不先做 NFC、大小写、换行或空白归一化;SHA256 对原文 UTF-8 字节计算。tokens 连续覆盖全文,拼接 text 必须逐字符等于原文,空白也有独立区间。空串合法;最多 100000 Unicode code point,拒绝孤立代理项。kind 为 word/space/punctuation;word 是本小样的可点 token 类别,也可能包含数字或 emoji,不保证是自然语言词条。
|
||||
|
||||
每个 token 提供 text、lemma、kind 与半开区间 [start,end):
|
||||
|
||||
| 字段后缀 | 单位与使用方 |
|
||||
|---|---|
|
||||
| cp | Unicode code point,Python 字符串索引;不是用户感知字形 |
|
||||
| utf8 | UTF-8 字节,可供 Go string 字节切片 |
|
||||
| utf16 | UTF-16 code unit,JavaScript String.slice / DOM 文本位置 |
|
||||
|
||||
例:原文 `A🙂é`,emoji 的 cp=[1,2)、utf8=[1,5)、utf16=[1,3);后面的 e 加组合重音共两个 code point,cp=[2,4)、utf8=[5,8)、utf16=[3,5)。不能把这些单位混用,也不能把组合字符或 ZWJ 序列的 code point 数当作可见字符数。浏览器先逐 token 校验 UTF-16 切片并检查完整重建再展示;textarea 会按浏览器规范将换行转为 LF,所以 HTTP 契约保证收到的原文,不承诺还原剪贴板进入 textarea 前的 CRLF。服务端 CRLF 原文测试单独覆盖。
|
||||
|
||||
POST /lookup 接收 {surface,lemma?}。查词键单独 casefold/NFC/弯撇号转 ASCII,不改变原文位置;先精确查询 surface,再尝试调用方提供的 lemma。结果 status 为 exact、lemma、not_found 或 resource_missing,含 matched_form 和最多 12 条 entries(lemma/pos/definition/examples)。`dog` 直接命中;点击 `went` 可用上下文 lemma `go` 回退;手动只输入 `went` 不猜词性而返回未找到。词典缺失和模型缺失分别标识 wordnet/model,不能伪装成查无结果。
|
||||
|
||||
实验查询无用户状态、写入或缓存私人输入;重复查询确定性返回,原文哈希可检测文本版本变化,但尚未定义生产 token ID、任务幂等或个人词语合并规则。lemma 不等于学习状态身份,禁止自动合并原词/词元。WordNet 仅英英释义,按 n/v/a/r 与原生 sense 顺序截取,不做上下文义项排序、翻译或发音;`The leaves fell.` 的 leaves 实测被模型错误还原为 leave,此限制保留供后续用户选择/修正方案参考。
|
||||
|
||||
@@ -2,8 +2,8 @@
|
||||
generated: true (请先修改 Gitea Wiki,禁止直接编辑本文件)
|
||||
wiki_page: Local-Development-and-Verification
|
||||
wiki_url: https://git.ilapage.cn/OPC/lexgo/wiki/Local-Development-and-Verification.-
|
||||
wiki_revision: 8857cf2545a3de6f1efa07d88b920a8301080e56
|
||||
synchronized_at: 2026-09-10T12:43:25Z
|
||||
wiki_revision: bf3483ec95bb73d522776d89dd2c20f9d117adb5
|
||||
synchronized_at: 2026-09-10T12:58:36Z
|
||||
<!-- gitea-wiki-mirror:end -->
|
||||
|
||||
# 本地开发与验证
|
||||
@@ -192,3 +192,36 @@ supervisor 直接管理编译后的 Go 进程,运行时不调用 Python。数
|
||||
|
||||
|
||||
#18 用户验收:2026-09-10T20:40:04+08:00 用户确认日志通过验收(工单评论 7591),包含此前待人工检查的交互。未重新运行自动化测试,未更改其历史结果,未合并 PR 或发布生产。
|
||||
|
||||
|
||||
## #3 英语离线验证入口与复现
|
||||
|
||||
从仓库根目录执行(uv 与 Node 已安装,不能使用本机默认 Python 3.8):
|
||||
|
||||
```powershell
|
||||
uv venv --python 3.12.12 .local/nlp-venv
|
||||
uv pip install --python .local/nlp-venv/Scripts/python.exe -r spikes/english/requirements.lock
|
||||
.local/nlp-venv/Scripts/python.exe spikes/english/setup_resources.py
|
||||
uv pip install --python .local/nlp-venv/Scripts/python.exe --no-deps .local/nlp-resources/en_core_web_sm-3.8.0-py3-none-any.whl
|
||||
.local/nlp-venv/Scripts/python.exe -m unittest discover -s spikes/english -v
|
||||
node --test spikes/english/view.test.mjs
|
||||
.local/nlp-venv/Scripts/python.exe spikes/english/benchmark.py
|
||||
.local/nlp-venv/Scripts/python.exe spikes/english/app.py
|
||||
```
|
||||
|
||||
打开 http://127.0.0.1:5183/,默认虚构样例,分析后点击 went 应出现 go 与“按原形查询”;dog 直接命中,zzzxqvfiction 未找到。`--resources .local/absent-resources` 可验证词典缺失,`--port` 可更换临时端口。模型缺失、词典缺失、非法输入及内部错误不输出路径/正文。仅 loopback,Host/Origin 校验,禁跨域、无缓存、无访问日志、连接读超时 10 秒。未配置 supervisor;停止该临时进程即可回退,既有服务和数据不变。
|
||||
|
||||
首次准备需要联网,失败可重跑;资源文件通过固定 SHA256 校验后使用。模型 3.8.0 MIT,WordNet 3.0 ZIP 完整保留 LICENSE/版权/免责声明,spaCy MIT、NLTK Apache-2.0。固定资源 URL 和摘要见 spikes/english/resources.json,原始许可与来源见该目录 README。词典为英英格式,不是中文翻译库。
|
||||
|
||||
2026-09-10 实测:11 项 Python 测试和 2 项 JavaScript 测试通过。真实模型/词典测试及 benchmark 禁止 socket connect,验证运行期无在线翻译依赖;不是整机断网测试。浏览器已实测展示分词、点击 went→go、手动查询无结果。外部 spaCy/Click 有一条 DeprecationWarning,未影响测试结果。Windows 10 19044,Intel Family 6 Model 140、8 逻辑核,Python 3.12.12;完整环境、UTC 时间和样本保存在 benchmark-result.json。
|
||||
|
||||
| 测量 | 本次样本 |
|
||||
|---|---|
|
||||
| 冷进程 Engine 加载(含 import,文件系统缓存可能已热) | 2355 ms |
|
||||
| 首次分析 / 首次 dog 查询 | 见 benchmark-result.json(各 1 次) |
|
||||
| 100000 code point(106095 UTF-8 字节),spaCy 3 次 | 中位 1433 ms,约 6.98 万 cp/s |
|
||||
| 同文正则+WordNet morphology,3 次 | 中位约 336 ms;未包含三位置转换,非完全等价负载 |
|
||||
| dog 查询,热进程 100 次 | 中位 0.0149 ms,p95 0.023 ms |
|
||||
| 12 个显式 lemma 样例 | spaCy 11/12、基线 6/12;保留 leaves 错误 |
|
||||
|
||||
推荐 Python NLP,但该样本不代表一般准确率、生产并发能力或延迟保证。尚未验证正式 Go/Python 调用、长任务持久化、移动端划词(#4)、英汉词典及生产部署。
|
||||
|
||||
+5
-2
@@ -2,8 +2,8 @@
|
||||
generated: true (请先修改 Gitea Wiki,禁止直接编辑本文件)
|
||||
wiki_page: Home
|
||||
wiki_url: https://git.ilapage.cn/OPC/lexgo/wiki/Home
|
||||
wiki_revision: d113bfc4d8d33feced5b84f098a5c50ba51df0ba
|
||||
synchronized_at: 2026-09-10T12:43:18Z
|
||||
wiki_revision: c5a40a06bd288faab9c0a3cbebb3bc2713bc13e7
|
||||
synchronized_at: 2026-09-10T12:58:28Z
|
||||
<!-- gitea-wiki-mirror:end -->
|
||||
|
||||
# LexGo 文档入口
|
||||
@@ -57,3 +57,6 @@ Quant-UX 原型 v1 已通过用户验收。[桌面预览](https://qux.ilapage.cn
|
||||
|
||||
|
||||
日志审计 #18 于 2026-09-10T20:40:04+08:00 获用户验收并关闭;#16 已更新完成索引。下一阶段 #3/#4 尚未开始。
|
||||
|
||||
|
||||
英语分词/原文定位/本地词典 #3 已完成独立可运行小样,待用户验收。代码与复现命令位于 spikes/english;临时入口 http://127.0.0.1:5183/。推荐 Python spaCy 英语模型与 WordNet 3.0 的离线组合;尚未接入正式学习端,#4 划词验证仍未开始。
|
||||
|
||||
@@ -0,0 +1,31 @@
|
||||
# #3 英语技术验证
|
||||
|
||||
独立小样,不是学习端正式功能;不连接 MySQL、不读取账号、不保存输入。仅本机、单进程串行运行,默认英语。长期契约与实测结论见 [开发验证 Wiki](https://git.ilapage.cn/OPC/lexgo/wiki/Local-Development-and-Verification)。
|
||||
|
||||
从仓库根目录运行(Windows PowerShell,需要 uv、Node):
|
||||
|
||||
```powershell
|
||||
uv venv --python 3.12.12 .local/nlp-venv
|
||||
uv pip install --python .local/nlp-venv/Scripts/python.exe -r spikes/english/requirements.lock
|
||||
.local/nlp-venv/Scripts/python.exe spikes/english/setup_resources.py
|
||||
uv pip install --python .local/nlp-venv/Scripts/python.exe --no-deps .local/nlp-resources/en_core_web_sm-3.8.0-py3-none-any.whl
|
||||
.local/nlp-venv/Scripts/python.exe -m unittest discover -s spikes/english -v
|
||||
node --test spikes/english/view.test.mjs
|
||||
.local/nlp-venv/Scripts/python.exe spikes/english/benchmark.py
|
||||
.local/nlp-venv/Scripts/python.exe spikes/english/app.py
|
||||
```
|
||||
|
||||
打开 <http://127.0.0.1:5183/>。点击“分析文本”后点单词;`went` 应以 `go` 查询。手动查询仅查输入形式,不猜测词性:`dog` 直接命中,`went` 无结果。关闭进程即停止小样;端口占用时使用 `--port 5185`。用 `--resources .local/absent-resources` 启动可验证词典缺失,模型与词典独立加载。
|
||||
|
||||
准备依赖和资源时需要联网;安装完成后运行不依赖在线翻译或下载服务。`test_engine.py` 与 `benchmark.py` 禁止 socket connect,用真实模型与词典验证离线运行。网络下载失败可重新执行准备命令,已有资源先校验再复用。默认系统 Python 3.8 不适用,命令必须使用上述独立环境。
|
||||
|
||||
资源版本、固定下载 URL 和 SHA256 见 `resources.json`,Python 依赖固定于 `requirements.lock`。大文件只保存在忽略的 `.local/nlp-resources`。
|
||||
|
||||
许可与来源:
|
||||
|
||||
- [spaCy 3.8.7](https://pypi.org/pypi/spacy/3.8.7/json) 与 [en_core_web_sm 3.8.0](https://github.com/explosion/spacy-models/releases/tag/en_core_web_sm-3.8.0):MIT;安装包保留其许可证。模型的 POS/lemma 组件保留,parser/NER 在本小样中停用;不输出句界。
|
||||
- [NLTK](https://github.com/nltk/nltk/blob/3.9.2/LICENSE.txt):Apache-2.0;只读取本地词典文件,无隐式 downloader。
|
||||
- [Princeton WordNet 3.0](https://wordnet.princeton.edu/license-and-commercial-use):WordNet 3.0 许可证;下载 ZIP 完整保留 `wordnet/LICENSE`、版权及免责声明。英英释义,不提供中文翻译。再分发必须保留许可声明。
|
||||
- [WordNet 原生格式](https://wordnet.princeton.edu/documentation/wndb5wn) 是 index/data/exception 文件,本小样直接读取 ZIP 中原始文件,不使用 go-admin 的系统枚举字典。
|
||||
|
||||
`benchmark-result.json` 是本机虚构语料的测量样本,不代表一般准确率或生产性能承诺。
|
||||
@@ -0,0 +1,69 @@
|
||||
import {validateTokens,latestOnly} from './view.mjs'
|
||||
const source=document.querySelector('#source'), reading=document.querySelector('#reading'), definition=document.querySelector('#definition'), notice=document.querySelector('#notice'), position=document.querySelector('#position'), query=document.querySelector('#query'), analyzeButton=document.querySelector('#analyze')
|
||||
const analysis=latestOnly(), lookup=latestOnly()
|
||||
const labels={exact:'直接命中',lemma:'按原形查询',not_found:'未找到释义',resource_missing:'本地资源缺失',invalid_input:'输入格式无效或文本超过限制',internal_error:'暂时无法处理,请重试',too_large:'文本超过限制'}
|
||||
async function post(path,payload) {
|
||||
const response=await fetch(path,{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify(payload)})
|
||||
const data=await response.json()
|
||||
if(!response.ok) throw new Error(labels[data.status] || '请求失败,请重试')
|
||||
return data
|
||||
}
|
||||
function element(tag,text) {const node=document.createElement(tag);node.textContent=text;return node}
|
||||
async function showWord(surface,lemma='',token=null) {
|
||||
const version=lookup.next()
|
||||
query.value=surface
|
||||
definition.replaceChildren(element('p','查询中…'))
|
||||
position.textContent=token ? JSON.stringify({text:token.text,lemma:token.lemma,code_point:[token.start_cp,token.end_cp],utf8:[token.start_utf8,token.end_utf8],utf16:[token.start_utf16,token.end_utf16]},null,2) : '手动查询,无原文位置'
|
||||
try {
|
||||
const data=await post('/lookup',{surface,lemma})
|
||||
if(!lookup.current(version)) return
|
||||
definition.replaceChildren(element('h3',data.matched_form || surface),element('small',labels[data.status] || data.status))
|
||||
if(data.entries.length) {
|
||||
const list=document.createElement('ol')
|
||||
for(const entry of data.entries) {
|
||||
const item=element('li',entry.definition)
|
||||
item.prepend(element('small',entry.pos+' · '))
|
||||
for(const example of entry.examples) item.append(element('p',example))
|
||||
list.append(item)
|
||||
}
|
||||
definition.append(list)
|
||||
}
|
||||
} catch(error) {if(lookup.current(version)) definition.textContent=error.message}
|
||||
}
|
||||
function resetReading() {
|
||||
lookup.next()
|
||||
reading.replaceChildren()
|
||||
definition.textContent='点击文中的单词。'
|
||||
position.textContent='尚未选择单词'
|
||||
}
|
||||
async function analyze() {
|
||||
const version=analysis.next(), text=source.value
|
||||
resetReading()
|
||||
analyzeButton.disabled=true
|
||||
notice.textContent='正在分析…'
|
||||
try {
|
||||
const result=await post('/analyze',{text})
|
||||
if(!analysis.current(version)) return
|
||||
validateTokens(text,result.tokens)
|
||||
const fragment=document.createDocumentFragment()
|
||||
for(const token of result.tokens) {
|
||||
if(token.kind==='word') {
|
||||
const button=element('button',token.text)
|
||||
button.type='button'
|
||||
button.addEventListener('click',()=>{
|
||||
reading.querySelector('.selected')?.classList.remove('selected')
|
||||
button.classList.add('selected')
|
||||
showWord(token.text,token.lemma,token)
|
||||
})
|
||||
fragment.append(button)
|
||||
} else fragment.append(document.createTextNode(token.text))
|
||||
}
|
||||
reading.replaceChildren(fragment)
|
||||
notice.textContent=text ? '已分析 · 点击单词查词' : '请输入阅读文本'
|
||||
} catch(error) {if(analysis.current(version)) notice.textContent=error.message}
|
||||
finally {if(analysis.current(version)) analyzeButton.disabled=false}
|
||||
}
|
||||
source.addEventListener('input',()=>{analysis.next();resetReading();analyzeButton.disabled=false;notice.textContent='文本已修改,请重新分析'})
|
||||
analyzeButton.addEventListener('click',analyze)
|
||||
document.querySelector('#lookup-form').addEventListener('submit',event=>{event.preventDefault();showWord(query.value)})
|
||||
analyze()
|
||||
@@ -0,0 +1,117 @@
|
||||
"""Loopback-only, ephemeral English experiment. Not a production API."""
|
||||
import argparse
|
||||
import json
|
||||
from http.server import BaseHTTPRequestHandler, HTTPServer
|
||||
from pathlib import Path
|
||||
|
||||
STATIC = Path(__file__).parent
|
||||
ROOT = STATIC.parent.parent
|
||||
|
||||
|
||||
def create_server(engine, port=5183):
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
def setup(self):
|
||||
super().setup()
|
||||
self.connection.settimeout(10)
|
||||
|
||||
def log_message(self, *_args):
|
||||
pass # Do not retain input or query text in access logs.
|
||||
|
||||
def reply(self, code, body, mime='application/json; charset=utf-8'):
|
||||
if not isinstance(body, bytes):
|
||||
body = json.dumps(body, ensure_ascii=True).encode('utf-8')
|
||||
self.send_response(code)
|
||||
self.send_header('Content-Type', mime)
|
||||
self.send_header('Content-Length', str(len(body)))
|
||||
self.send_header('Cache-Control', 'no-store')
|
||||
self.send_header('X-Content-Type-Options', 'nosniff')
|
||||
self.send_header('Content-Security-Policy', "default-src 'self'; style-src 'self'; script-src 'self'; connect-src 'self'; frame-ancestors 'none'; base-uri 'none'")
|
||||
self.end_headers()
|
||||
self.wfile.write(body)
|
||||
|
||||
def local_request(self):
|
||||
host = self.headers.get('Host')
|
||||
allowed = {f'127.0.0.1:{self.server.server_port}', f'localhost:{self.server.server_port}'}
|
||||
origin = self.headers.get('Origin')
|
||||
if host not in allowed or (origin is not None and origin not in {'http://' + h for h in allowed}):
|
||||
self.reply(403, {'status': 'forbidden'})
|
||||
return False
|
||||
return True
|
||||
|
||||
def do_GET(self):
|
||||
if not self.local_request():
|
||||
return
|
||||
files = {'/': ('index.html', 'text/html'), '/app.mjs': ('app.mjs', 'text/javascript'), '/view.mjs': ('view.mjs', 'text/javascript'), '/style.css': ('style.css', 'text/css')}
|
||||
if self.path not in files:
|
||||
self.reply(404, {'status': 'not_found'})
|
||||
return
|
||||
name, mime = files[self.path]
|
||||
self.reply(200, (STATIC / name).read_bytes(), mime + '; charset=utf-8')
|
||||
|
||||
def do_POST(self):
|
||||
if not self.local_request():
|
||||
return
|
||||
if self.path not in {'/analyze', '/lookup'}:
|
||||
self.reply(404, {'status': 'not_found'})
|
||||
return
|
||||
try:
|
||||
length = int(self.headers.get('Content-Length', '0'))
|
||||
except ValueError:
|
||||
self.reply(400, {'status': 'invalid_input'})
|
||||
return
|
||||
if length > 1_000_000:
|
||||
self.reply(413, {'status': 'too_large'})
|
||||
return
|
||||
if length <= 0:
|
||||
self.reply(400, {'status': 'invalid_input'})
|
||||
return
|
||||
if self.headers.get_content_type() != 'application/json':
|
||||
self.reply(415, {'status': 'invalid_content_type'})
|
||||
return
|
||||
try:
|
||||
payload = json.loads(self.rfile.read(length))
|
||||
if not isinstance(payload, dict):
|
||||
raise ValueError('object required')
|
||||
key = 'text' if self.path == '/analyze' else 'surface'
|
||||
if not isinstance(payload.get(key), str):
|
||||
raise ValueError('string required')
|
||||
if self.path == '/analyze':
|
||||
result = engine.analyze(payload['text'])
|
||||
else:
|
||||
lemma = payload.get('lemma', '')
|
||||
if not isinstance(lemma, str) or len(lemma) > 200 or len(payload['surface']) > 200:
|
||||
raise ValueError('invalid query')
|
||||
result = engine.lookup(payload['surface'], lemma)
|
||||
self.reply(200, result)
|
||||
except (ValueError, TypeError, UnicodeError):
|
||||
self.reply(400, {'status': 'invalid_input'})
|
||||
except Exception as exc:
|
||||
# ResourceMissing is deliberately exposed without filesystem paths.
|
||||
from engine import ResourceMissing
|
||||
if isinstance(exc, ResourceMissing):
|
||||
self.reply(503, {'status': 'resource_missing', 'resource': exc.resource})
|
||||
else:
|
||||
self.reply(500, {'status': 'internal_error'})
|
||||
|
||||
server = HTTPServer(('127.0.0.1', port), Handler)
|
||||
return server
|
||||
|
||||
|
||||
def main():
|
||||
from engine import Engine
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument('--port', type=int, default=5183)
|
||||
parser.add_argument('--resources', type=Path, default=ROOT / '.local/nlp-resources')
|
||||
args = parser.parse_args()
|
||||
server = create_server(Engine(args.resources), args.port)
|
||||
print(f'LexGo English experiment: http://127.0.0.1:{server.server_port}/', flush=True)
|
||||
try:
|
||||
server.serve_forever()
|
||||
except KeyboardInterrupt:
|
||||
pass
|
||||
finally:
|
||||
server.server_close()
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1,181 @@
|
||||
{
|
||||
"measured_at_utc": "2026-09-10T12:55:12.016850+00:00",
|
||||
"environment": {
|
||||
"python": "3.12.12",
|
||||
"platform": "Windows-10-10.0.19044-SP0",
|
||||
"processor": "Intel64 Family 6 Model 140 Stepping 1, GenuineIntel",
|
||||
"logical_cpus": 8,
|
||||
"packages": {
|
||||
"spacy": "3.8.7",
|
||||
"nltk": "3.9.2",
|
||||
"en-core-web-sm": "3.8.0"
|
||||
}
|
||||
},
|
||||
"offline": true,
|
||||
"cold_engine_load_ms": 2355.1940000616014,
|
||||
"text_codepoints": 100000,
|
||||
"text_utf8_bytes": 106095,
|
||||
"quality": {
|
||||
"cases": [
|
||||
{
|
||||
"sentence": "She went home.",
|
||||
"surface": "went",
|
||||
"expected": "go",
|
||||
"spacy": "go",
|
||||
"baseline": "go",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": true
|
||||
},
|
||||
{
|
||||
"sentence": "The children ate apples.",
|
||||
"surface": "children",
|
||||
"expected": "child",
|
||||
"spacy": "child",
|
||||
"baseline": "child",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": true
|
||||
},
|
||||
{
|
||||
"sentence": "The children ate apples.",
|
||||
"surface": "ate",
|
||||
"expected": "eat",
|
||||
"spacy": "eat",
|
||||
"baseline": "ate",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": false
|
||||
},
|
||||
{
|
||||
"sentence": "The dogs ran quickly.",
|
||||
"surface": "dogs",
|
||||
"expected": "dog",
|
||||
"spacy": "dog",
|
||||
"baseline": "dog",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": true
|
||||
},
|
||||
{
|
||||
"sentence": "The dogs ran quickly.",
|
||||
"surface": "ran",
|
||||
"expected": "run",
|
||||
"spacy": "run",
|
||||
"baseline": "run",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": true
|
||||
},
|
||||
{
|
||||
"sentence": "I saw a bird.",
|
||||
"surface": "saw",
|
||||
"expected": "see",
|
||||
"spacy": "see",
|
||||
"baseline": "saw",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": false
|
||||
},
|
||||
{
|
||||
"sentence": "The saw is sharp.",
|
||||
"surface": "saw",
|
||||
"expected": "saw",
|
||||
"spacy": "saw",
|
||||
"baseline": "saw",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": true
|
||||
},
|
||||
{
|
||||
"sentence": "She leaves today.",
|
||||
"surface": "leaves",
|
||||
"expected": "leave",
|
||||
"spacy": "leave",
|
||||
"baseline": "leaf",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": false
|
||||
},
|
||||
{
|
||||
"sentence": "The leaves fell.",
|
||||
"surface": "leaves",
|
||||
"expected": "leaf",
|
||||
"spacy": "leave",
|
||||
"baseline": "leaf",
|
||||
"spacy_correct": false,
|
||||
"baseline_correct": true
|
||||
},
|
||||
{
|
||||
"sentence": "They are reading books.",
|
||||
"surface": "reading",
|
||||
"expected": "read",
|
||||
"spacy": "read",
|
||||
"baseline": "reading",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": false
|
||||
},
|
||||
{
|
||||
"sentence": "He was better yesterday.",
|
||||
"surface": "was",
|
||||
"expected": "be",
|
||||
"spacy": "be",
|
||||
"baseline": "wa",
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": false
|
||||
},
|
||||
{
|
||||
"sentence": "She can't go.",
|
||||
"surface": "n't",
|
||||
"expected": "not",
|
||||
"spacy": "not",
|
||||
"baseline": null,
|
||||
"spacy_correct": true,
|
||||
"baseline_correct": false
|
||||
}
|
||||
],
|
||||
"total": 12,
|
||||
"spacy_correct": 11,
|
||||
"baseline_correct": 6
|
||||
},
|
||||
"measurements": {
|
||||
"first_analysis": {
|
||||
"repeats": 1,
|
||||
"median_ms": 2.9011000879108906,
|
||||
"p95_ms": 2.9011000879108906
|
||||
},
|
||||
"first_lookup": {
|
||||
"repeats": 1,
|
||||
"median_ms": 78.619199921377,
|
||||
"p95_ms": 78.619199921377
|
||||
},
|
||||
"spacy_100k": {
|
||||
"repeats": 3,
|
||||
"median_ms": 1432.92090005707,
|
||||
"p95_ms": 1499.0439999382943,
|
||||
"codepoints_per_second": 69787.52281163407
|
||||
},
|
||||
"baseline_100k": {
|
||||
"repeats": 3,
|
||||
"median_ms": 336.1744999419898,
|
||||
"p95_ms": 343.62249996047467,
|
||||
"codepoints_per_second": 297464.56086721626
|
||||
},
|
||||
"lookup": {
|
||||
"dog": {
|
||||
"repeats": 100,
|
||||
"median_ms": 0.014899997040629387,
|
||||
"p95_ms": 0.022999942302703857
|
||||
},
|
||||
"went": {
|
||||
"repeats": 100,
|
||||
"median_ms": 0.021250045392662287,
|
||||
"p95_ms": 0.02929999027401209
|
||||
},
|
||||
"zzzxqvfiction": {
|
||||
"repeats": 100,
|
||||
"median_ms": 0.003600027412176132,
|
||||
"p95_ms": 0.004999921657145023
|
||||
}
|
||||
}
|
||||
},
|
||||
"limitations": [
|
||||
"Synthetic microbenchmark; 12 selected cases do not establish general accuracy.",
|
||||
"Baseline excludes UTF-8/UTF-16 conversion; spaCy timing includes full analyze contract.",
|
||||
"Repeated queries are warm-process; no claim about production concurrency.",
|
||||
"Cold engine load includes dependency imports but OS filesystem caches may be warm.",
|
||||
"POS and sense disambiguation are not provided by dictionary lookup."
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,112 @@
|
||||
"""Offline synthetic measurements; stdout is one JSON document, no text input log."""
|
||||
import argparse
|
||||
from datetime import datetime, timezone
|
||||
import importlib.metadata
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
from pathlib import Path
|
||||
import platform
|
||||
import re
|
||||
import statistics
|
||||
import time
|
||||
from unittest.mock import patch
|
||||
|
||||
from engine import Engine
|
||||
|
||||
|
||||
CASES = [
|
||||
('She went home.', 'went', 'go'),
|
||||
('The children ate apples.', 'children', 'child'),
|
||||
('The children ate apples.', 'ate', 'eat'),
|
||||
('The dogs ran quickly.', 'dogs', 'dog'),
|
||||
('The dogs ran quickly.', 'ran', 'run'),
|
||||
('I saw a bird.', 'saw', 'see'),
|
||||
('The saw is sharp.', 'saw', 'saw'),
|
||||
('She leaves today.', 'leaves', 'leave'),
|
||||
('The leaves fell.', 'leaves', 'leaf'),
|
||||
('They are reading books.', 'reading', 'read'),
|
||||
('He was better yesterday.', 'was', 'be'),
|
||||
("She can't go.", "n't", 'not'),
|
||||
]
|
||||
|
||||
|
||||
def summary(samples):
|
||||
values = sorted(samples)
|
||||
return {'repeats': len(values), 'median_ms': statistics.median(values) * 1000,
|
||||
'p95_ms': values[max(0, math.ceil(len(values) * .95) - 1)] * 1000}
|
||||
|
||||
|
||||
def measure(function, repeats):
|
||||
samples = []
|
||||
for _ in range(repeats):
|
||||
start = time.perf_counter()
|
||||
function()
|
||||
samples.append(time.perf_counter() - start)
|
||||
return summary(samples)
|
||||
|
||||
|
||||
def run(args):
|
||||
start = time.perf_counter()
|
||||
engine = Engine(args.resources)
|
||||
load_seconds = time.perf_counter() - start
|
||||
if engine.nlp is None or engine.wordnet is None:
|
||||
raise RuntimeError(f'Resources unavailable: model={engine.model_error}, wordnet={engine.wordnet_error}')
|
||||
measured = {}
|
||||
measured['first_analysis'] = measure(lambda: engine.analyze('She went home.'), 1)
|
||||
measured['first_lookup'] = measure(lambda: engine.lookup('dog'), 1)
|
||||
baseline_pattern = re.compile(r"\w+(?:['’]\w+)*|\s+|[^\w\s]", re.UNICODE)
|
||||
# Lower-cost baseline: regex spans and context-free WordNet morphology.
|
||||
def baseline(text):
|
||||
return [(match.group(), engine.wordnet.morphy(match.group().lower()) or match.group().lower(),
|
||||
match.start(), match.end()) for match in baseline_pattern.finditer(text)]
|
||||
|
||||
quality = []
|
||||
for sentence, surface, expected in CASES:
|
||||
contextual = next((token['lemma'] for token in engine.analyze(sentence)['tokens']
|
||||
if token['text'] == surface), None)
|
||||
simple = next((lemma for token, lemma, _, _ in baseline(sentence) if token == surface), None)
|
||||
quality.append(dict(sentence=sentence, surface=surface, expected=expected,
|
||||
spacy=contextual, baseline=simple,
|
||||
spacy_correct=contextual == expected, baseline_correct=simple == expected))
|
||||
seed = "She went home. The children ate apples. I saw a bird. The leaves fell. Café 😀 e\u0301\r\n"
|
||||
text = (seed * (100000 // len(seed) + 1))[:100000]
|
||||
for name, function in [('spacy', engine.analyze), ('baseline', baseline)]:
|
||||
metrics = measure(lambda: function(text), args.text_repeats)
|
||||
metrics['codepoints_per_second'] = len(text) / (metrics['median_ms'] / 1000)
|
||||
measured[name + '_100k'] = metrics
|
||||
queries = [('dog', ''), ('went', 'go'), ('zzzxqvfiction', '')]
|
||||
measured['lookup'] = {surface: measure(lambda: engine.lookup(surface, lemma), args.query_repeats)
|
||||
for surface, lemma in queries}
|
||||
return {
|
||||
'measured_at_utc': datetime.now(timezone.utc).isoformat(),
|
||||
'environment': {'python': platform.python_version(), 'platform': platform.platform(),
|
||||
'processor': platform.processor(), 'logical_cpus': os.cpu_count(),
|
||||
'packages': {name: importlib.metadata.version(name)
|
||||
for name in ('spacy', 'nltk', 'en-core-web-sm')}},
|
||||
'offline': True, 'cold_engine_load_ms': load_seconds * 1000,
|
||||
'text_codepoints': len(text), 'text_utf8_bytes': len(text.encode()),
|
||||
'quality': {'cases': quality, 'total': len(quality),
|
||||
'spacy_correct': sum(case['spacy_correct'] for case in quality),
|
||||
'baseline_correct': sum(case['baseline_correct'] for case in quality)},
|
||||
'measurements': measured,
|
||||
'limitations': [
|
||||
'Synthetic microbenchmark; 12 selected cases do not establish general accuracy.',
|
||||
'Baseline excludes UTF-8/UTF-16 conversion; spaCy timing includes full analyze contract.',
|
||||
'Repeated queries are warm-process; no claim about production concurrency.',
|
||||
'Cold engine load includes dependency imports but OS filesystem caches may be warm.',
|
||||
'POS and sense disambiguation are not provided by dictionary lookup.',
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument('--resources', type=Path, default=Path(__file__).resolve().parents[2] / '.local/nlp-resources')
|
||||
parser.add_argument('--text-repeats', type=int, default=3)
|
||||
parser.add_argument('--query-repeats', type=int, default=100)
|
||||
args = parser.parse_args()
|
||||
if args.text_repeats < 1 or args.query_repeats < 1:
|
||||
parser.error('repeat counts must be positive')
|
||||
with patch('socket.socket.connect', side_effect=AssertionError('network forbidden')):
|
||||
print(json.dumps(run(args), ensure_ascii=True, indent=2))
|
||||
@@ -0,0 +1,116 @@
|
||||
"""Isolated English experiment. Resources are local; no downloader is used."""
|
||||
import hashlib
|
||||
from pathlib import Path
|
||||
import unicodedata
|
||||
import warnings
|
||||
import zipfile
|
||||
|
||||
|
||||
class ResourceMissing(RuntimeError):
|
||||
def __init__(self, resource):
|
||||
self.resource = resource
|
||||
super().__init__(f'Local {resource} resource is unavailable')
|
||||
|
||||
|
||||
def validate_text(text):
|
||||
if not isinstance(text, str):
|
||||
raise TypeError('text must be a string')
|
||||
if len(text) > 100000:
|
||||
raise ValueError('text exceeds 100000 code points')
|
||||
if any(0xD800 <= ord(char) <= 0xDFFF for char in text):
|
||||
raise ValueError('text contains an unpaired surrogate')
|
||||
|
||||
|
||||
def lookup_form(text):
|
||||
validate_text(text)
|
||||
return unicodedata.normalize('NFC', text.casefold()).replace('’', "'").replace('‘', "'")
|
||||
|
||||
|
||||
class Engine:
|
||||
def __init__(self, resource_dir: Path, model_name='en_core_web_sm'):
|
||||
self.nlp = None
|
||||
self.wordnet = None
|
||||
self.model_error = None
|
||||
self.wordnet_error = None
|
||||
try:
|
||||
import spacy
|
||||
self.nlp = spacy.load(model_name, disable=['parser', 'ner'])
|
||||
except (ImportError, OSError, ValueError) as error:
|
||||
self.model_error = type(error).__name__
|
||||
try:
|
||||
from nltk.corpus.reader import WordNetCorpusReader
|
||||
from nltk.data import ZipFilePathPointer
|
||||
|
||||
class EnglishWordNet30Reader(WordNetCorpusReader):
|
||||
def map_wn(self, version='wordnet'):
|
||||
# NLTK's default cross-version OMW mapping loads a global
|
||||
# corpus. English-only WordNet 3.0 needs no such mapping.
|
||||
if self.get_version() != '3.0':
|
||||
raise ValueError('This experiment requires WordNet 3.0')
|
||||
return None
|
||||
|
||||
root = ZipFilePathPointer(str(Path(resource_dir) / 'wordnet.zip'), 'wordnet/')
|
||||
with warnings.catch_warnings():
|
||||
warnings.filterwarnings('ignore', message='The multilingual functions are not available with this Wordnet version', category=UserWarning)
|
||||
self.wordnet = EnglishWordNet30Reader(root, None)
|
||||
except (ImportError, OSError, LookupError, ValueError, zipfile.BadZipFile) as error:
|
||||
self.wordnet_error = type(error).__name__
|
||||
|
||||
def analyze(self, text):
|
||||
validate_text(text)
|
||||
if self.nlp is None:
|
||||
raise ResourceMissing('model')
|
||||
# Prefix tables make conversion linear even for long Unicode documents.
|
||||
utf8 = [0]
|
||||
utf16 = [0]
|
||||
for char in text:
|
||||
utf8.append(utf8[-1] + len(char.encode('utf-8')))
|
||||
utf16.append(utf16[-1] + (2 if ord(char) > 0xFFFF else 1))
|
||||
tokens = []
|
||||
|
||||
def append(start, end, lemma, kind):
|
||||
tokens.append(dict(text=text[start:end], lemma=lemma, kind=kind,
|
||||
start_cp=start, end_cp=end,
|
||||
start_utf8=utf8[start], end_utf8=utf8[end],
|
||||
start_utf16=utf16[start], end_utf16=utf16[end]))
|
||||
|
||||
cursor = 0
|
||||
for token in self.nlp(text):
|
||||
if token.idx > cursor:
|
||||
append(cursor, token.idx, '', 'space')
|
||||
end = token.idx + len(token.text)
|
||||
kind = 'space' if token.is_space else 'punctuation' if token.is_punct else 'word'
|
||||
append(token.idx, end, token.lemma_ if kind == 'word' else '', kind)
|
||||
cursor = end
|
||||
if cursor < len(text):
|
||||
append(cursor, len(text), '', 'space')
|
||||
return dict(status='ok', contract_version='english-spike-v1', original_text=text,
|
||||
text_sha256=hashlib.sha256(text.encode('utf-8')).hexdigest(), tokens=tokens)
|
||||
|
||||
def _exact_entries(self, form):
|
||||
# NLTK 3.9.2 synsets() applies morphy even with check_exceptions=False.
|
||||
# Read its loaded index directly to keep exact and explicit lemma distinct.
|
||||
index = self.wordnet._lemma_pos_offset_map.get(form, {})
|
||||
entries = []
|
||||
for pos in ('n', 'v', 'a', 'r'):
|
||||
for offset in index.get(pos, []):
|
||||
synset = self.wordnet.synset_from_pos_and_offset(pos, offset)
|
||||
entries.append(dict(lemma=form, pos=synset.pos(),
|
||||
definition=synset.definition(), examples=synset.examples()))
|
||||
if len(entries) == 12:
|
||||
return entries
|
||||
return entries
|
||||
|
||||
def lookup(self, surface, lemma=''):
|
||||
form = lookup_form(surface)
|
||||
fallback = lookup_form(lemma)
|
||||
result = dict(query=surface, matched_form=None, entries=[])
|
||||
if self.wordnet is None:
|
||||
return dict(result, status='resource_missing', resource='wordnet')
|
||||
for candidate, status in ((form, 'exact'), (fallback, 'lemma')):
|
||||
if not candidate:
|
||||
continue
|
||||
entries = self._exact_entries(candidate)
|
||||
if entries:
|
||||
return dict(result, status=status, matched_form=candidate, entries=entries)
|
||||
return dict(result, status='not_found')
|
||||
@@ -0,0 +1,12 @@
|
||||
<!doctype html>
|
||||
<html lang="zh-CN">
|
||||
<head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>LexGo · 英语验证</title><link rel="stylesheet" href="/style.css"></head>
|
||||
<body>
|
||||
<header><a class="brand" href="/">LexGo<span>.</span></a><span class="badge">英语 · 技术验证</span></header>
|
||||
<main>
|
||||
<section class="input-section"><label for="source">阅读文本</label><textarea id="source" spellcheck="false" maxlength="100000">Mira’s well-known dogs went home. She can't wait.
|
||||
The children were running beside a café. 🙂 Café!</textarea><div class="actions"><span>仅本机处理 · 不保存</span><button id="analyze">分析文本</button></div></section>
|
||||
<p id="notice" role="status" aria-live="polite"></p>
|
||||
<div class="workspace"><section class="paper" aria-label="阅读结果"><h1>阅读</h1><div id="reading"></div></section><aside aria-label="词典"><div class="dictionary-header"><h2>本地词典</h2><span>WordNet 3.0</span></div><form id="lookup-form"><label class="sr-only" for="query">查询单词</label><input id="query" maxlength="200" placeholder="输入英语单词" autocomplete="off"><button>查询</button></form><div id="definition" aria-live="polite">点击文中的单词。</div><details><summary>原文位置</summary><pre id="position">尚未选择单词</pre></details></aside></div>
|
||||
</main><footer>独立验证小样 · 英语释义</footer><script type="module" src="/app.mjs"></script>
|
||||
</body></html>
|
||||
@@ -0,0 +1,47 @@
|
||||
annotated-doc==0.0.5
|
||||
annotated-types==0.8.0
|
||||
blis==1.3.3
|
||||
catalogue==2.0.10
|
||||
certifi==2026.7.22
|
||||
charset-normalizer==3.5.1
|
||||
click==8.5.0
|
||||
cloudpathlib==0.25.0
|
||||
cloudpickle==3.1.2
|
||||
colorama==0.4.6
|
||||
confection==0.1.5
|
||||
cymem==2.0.13
|
||||
idna==3.19
|
||||
jinja2==3.1.6
|
||||
joblib==1.6.0
|
||||
langcodes==3.5.1
|
||||
markdown-it-py==4.2.0
|
||||
markupsafe==3.0.3
|
||||
mdurl==0.1.2
|
||||
murmurhash==1.0.15
|
||||
nltk==3.9.2
|
||||
numpy==2.5.3
|
||||
packaging==26.3
|
||||
preshed==3.0.13
|
||||
pydantic==2.13.5
|
||||
pydantic-core==2.46.5
|
||||
pygments==2.21.0
|
||||
regex==2026.9.10
|
||||
requests==2.34.2
|
||||
rich==15.0.0
|
||||
setuptools==84.0.0
|
||||
shellingham==1.5.4
|
||||
smart-open==7.7.1
|
||||
spacy==3.8.7
|
||||
spacy-legacy==3.0.12
|
||||
spacy-loggers==1.0.5
|
||||
srsly==2.5.3
|
||||
thinc==8.3.11
|
||||
tqdm==4.70.0
|
||||
typer==0.27.2
|
||||
typer-slim==0.24.0
|
||||
typing-extensions==4.16.0
|
||||
typing-inspection==0.4.4
|
||||
urllib3==2.7.0
|
||||
wasabi==1.1.3
|
||||
weasel==0.4.3
|
||||
wrapt==2.4.0
|
||||
@@ -0,0 +1,21 @@
|
||||
{
|
||||
"python": "3.12.12",
|
||||
"spacy": "3.8.7",
|
||||
"nltk": "3.9.2",
|
||||
"model": {
|
||||
"name": "en_core_web_sm",
|
||||
"version": "3.8.0",
|
||||
"file": "en_core_web_sm-3.8.0-py3-none-any.whl",
|
||||
"url": "https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl",
|
||||
"sha256": "1932429db727d4bff3deed6b34cfc05df17794f4a52eeb26cf8928f7c1a0fb85",
|
||||
"license": "MIT"
|
||||
},
|
||||
"dictionary": {
|
||||
"name": "Princeton WordNet",
|
||||
"version": "3.0",
|
||||
"file": "wordnet.zip",
|
||||
"url": "https://raw.githubusercontent.com/nltk/nltk_data/96f9b3252457a2b97e52aec64c3dfceeb5c312d5/packages/corpora/wordnet.zip",
|
||||
"sha256": "cbda5ea6eef7f36a97a43d4a75f85e07fccbb4f23657d27b4ccbc93e2646ab59",
|
||||
"license": "WordNet 3.0 License (included in ZIP: wordnet/LICENSE)"
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,30 @@
|
||||
"""Explicit network-only preparation; runtime never calls this script."""
|
||||
import hashlib
|
||||
import json
|
||||
from pathlib import Path
|
||||
import urllib.request
|
||||
|
||||
HERE = Path(__file__).resolve().parent
|
||||
DEST = HERE.parent.parent / '.local/nlp-resources'
|
||||
|
||||
|
||||
def main():
|
||||
manifest = json.loads((HERE / 'resources.json').read_text(encoding='utf-8'))
|
||||
DEST.mkdir(parents=True, exist_ok=True)
|
||||
for key in ('model', 'dictionary'):
|
||||
item = manifest[key]
|
||||
target = DEST / item['file']
|
||||
if target.exists() and hashlib.sha256(target.read_bytes()).hexdigest() == item['sha256']:
|
||||
print(key + ': checksum verified', flush=True)
|
||||
continue
|
||||
req = urllib.request.Request(item['url'], headers={'User-Agent': 'LexGo-English-Spike/1'})
|
||||
with urllib.request.urlopen(req, timeout=120) as response:
|
||||
data = response.read(32 * 1024 * 1024 + 1)
|
||||
if hashlib.sha256(data).hexdigest() != item['sha256']:
|
||||
raise RuntimeError(key + ': checksum mismatch; resource not installed')
|
||||
target.write_bytes(data)
|
||||
print(key + ': downloaded and checksum verified', flush=True)
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
@@ -0,0 +1 @@
|
||||
:root{color:#253b32;background:#f5f4ef;font-family:"Segoe UI","Microsoft YaHei",sans-serif;font-synthesis:none}*{box-sizing:border-box}body{margin:0}header{height:78px;border-bottom:1px solid #dfe3da;padding:0 5%;display:flex;align-items:center;justify-content:space-between;background:#fff}.brand{font-size:29px;font-weight:700;text-decoration:none;color:inherit;letter-spacing:-1px}.brand span{color:#3b8060}.badge{font-size:13px;color:#66776a}main{max-width:1260px;margin:35px auto;padding:0 28px}.input-section{border-bottom:1px solid #d5ddd4;padding-bottom:25px}label{display:block;font-weight:600;margin-bottom:12px}textarea{display:block;width:100%;min-height:135px;resize:vertical;border:1px solid #cbd5cb;border-radius:8px;background:#fff;padding:16px;color:#263b30;font:18px/1.7 Georgia,serif}textarea:focus,input:focus,button:focus-visible{outline:2px solid #538967;outline-offset:3px}.actions{display:flex;justify-content:space-between;align-items:center;margin-top:12px}.actions span,footer{font-size:12px;color:#738075}button{background:#2c6347;color:white;border:0;border-radius:5px;padding:10px 20px;cursor:pointer;font:inherit}button:disabled{opacity:.55;cursor:wait}#notice{font-size:14px;min-height:20px}.workspace{display:grid;grid-template-columns:minmax(0,1fr) 330px;gap:25px}.paper,aside{background:#fff;border:1px solid #e0e5dc;border-radius:8px}.paper{padding:28px 32px;min-height:340px}h1{font-size:13px;letter-spacing:2px;color:#6f7e72;margin:0 0 26px}#reading{font:23px/1.95 Georgia,"Times New Roman",serif;white-space:pre-wrap;overflow-wrap:anywhere}#reading button{font:inherit;color:inherit;padding:0;border-radius:2px;background:transparent;text-align:left}#reading button:hover,#reading button.selected{background:#e1edcf;box-shadow:0 2px #658447}aside{padding:24px}.dictionary-header{display:flex;justify-content:space-between;align-items:center;margin-bottom:20px}.dictionary-header h2{font-size:17px;margin:0}.dictionary-header span{font-size:11px;color:#7e887f}form{display:flex;gap:7px;margin-bottom:22px}input{min-width:0;width:100%;padding:9px;border:1px solid #cbd5cb;border-radius:4px;font:inherit}form button{padding:9px 12px;white-space:nowrap}#definition{font-size:14px;line-height:1.7;overflow-wrap:anywhere}#definition h3{font:27px Georgia,serif;margin:0 0 8px}#definition ol{padding-left:21px}#definition li{margin-bottom:13px}#definition small{color:#637567}details{margin-top:25px;border-top:1px solid #e1e6de;padding-top:15px;color:#7a847d;font-size:12px}summary{cursor:pointer}pre{white-space:pre-wrap;overflow-wrap:anywhere;font-size:11px}footer{text-align:center;padding:35px}.sr-only{position:absolute;width:1px;height:1px;overflow:hidden;clip-path:inset(50%)}@media(max-width:750px){main{padding:0 16px;margin-top:20px}.workspace{grid-template-columns:1fr}.paper{padding:24px;min-height:230px}#reading{font-size:21px}.actions span{font-size:11px}}
|
||||
@@ -0,0 +1,72 @@
|
||||
import http.client
|
||||
import json
|
||||
import threading
|
||||
import unittest
|
||||
from unittest.mock import Mock
|
||||
|
||||
from app import create_server
|
||||
|
||||
|
||||
class HTTPTests(unittest.TestCase):
|
||||
def setUp(self):
|
||||
self.engine = Mock()
|
||||
self.engine.analyze.return_value = {'status': 'ok', 'tokens': []}
|
||||
self.engine.lookup.return_value = {'status': 'not_found', 'entries': []}
|
||||
self.server = create_server(self.engine, 0)
|
||||
self.thread = threading.Thread(target=self.server.serve_forever, daemon=True)
|
||||
self.thread.start()
|
||||
self.port = self.server.server_port
|
||||
|
||||
def tearDown(self):
|
||||
self.server.shutdown()
|
||||
self.server.server_close()
|
||||
self.thread.join()
|
||||
|
||||
def call(self, method, path, body=None, headers=None):
|
||||
c = http.client.HTTPConnection('127.0.0.1', self.port, timeout=3)
|
||||
c.request(method, path, body, headers or {})
|
||||
r = c.getresponse()
|
||||
result = r.status, r.read(), dict(r.getheaders())
|
||||
c.close()
|
||||
return result
|
||||
|
||||
def test_local_page_and_no_arbitrary_file_access(self):
|
||||
status, body, headers = self.call('GET', '/')
|
||||
self.assertEqual(status, 200)
|
||||
self.assertIn(b'LexGo', body)
|
||||
self.assertIn('Content-Security-Policy', headers)
|
||||
self.assertEqual(self.call('GET', '/../../.env.local')[0], 404)
|
||||
|
||||
def test_json_analyze_and_lookup(self):
|
||||
self.assertEqual(self.call('POST', '/analyze', json.dumps({'text': 'Hello'}), {'Content-Type': 'application/json'})[0], 200)
|
||||
self.engine.analyze.assert_called_once_with('Hello')
|
||||
self.assertEqual(self.call('POST', '/lookup', json.dumps({'surface': 'went', 'lemma': 'go'}), {'Content-Type': 'application/json'})[0], 200)
|
||||
self.engine.lookup.assert_called_once_with('went', 'go')
|
||||
|
||||
def test_reject_cross_origin_and_rebinding(self):
|
||||
for headers in ({'Origin': 'https://evil.example'}, {'Host': 'evil.example'}):
|
||||
self.assertEqual(self.call('POST', '/analyze', '{}', headers)[0], 403)
|
||||
self.engine.analyze.assert_not_called()
|
||||
|
||||
def test_invalid_payload_and_size(self):
|
||||
for payload in ('[]', '{}', '{', '{"text":42}'):
|
||||
self.assertEqual(self.call('POST', '/analyze', payload, {'Content-Type': 'application/json'})[0], 400)
|
||||
self.assertEqual(self.call('POST', '/analyze', '{}', {'Content-Type': 'text/plain'})[0], 415)
|
||||
self.assertEqual(self.call('POST', '/analyze', '{}', {'Content-Type': 'application/json', 'Content-Length': '1000001'})[0], 413)
|
||||
|
||||
def test_internal_errors_do_not_echo_input(self):
|
||||
self.engine.analyze.side_effect = RuntimeError('private sample')
|
||||
status, body, _ = self.call('POST', '/analyze', '{"text":"x"}', {'Content-Type': 'application/json'})
|
||||
self.assertEqual(status, 500)
|
||||
self.assertNotIn(b'private sample', body)
|
||||
|
||||
def test_missing_model_is_distinct_from_invalid_input(self):
|
||||
from engine import ResourceMissing
|
||||
self.engine.analyze.side_effect = ResourceMissing('model')
|
||||
status, body, _ = self.call('POST', '/analyze', '{"text":"x"}', {'Content-Type': 'application/json'})
|
||||
self.assertEqual(status, 503)
|
||||
self.assertEqual(json.loads(body), {'status': 'resource_missing', 'resource': 'model'})
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
unittest.main()
|
||||
@@ -0,0 +1,92 @@
|
||||
"""Run with the isolated Python: -m unittest discover -s spikes/english -v."""
|
||||
import hashlib
|
||||
from pathlib import Path
|
||||
import tempfile
|
||||
import unittest
|
||||
from unittest.mock import patch
|
||||
|
||||
try:
|
||||
from engine import Engine, ResourceMissing
|
||||
except ImportError:
|
||||
Engine = None
|
||||
|
||||
RESOURCES = Path(__file__).resolve().parents[2] / '.local/nlp-resources'
|
||||
|
||||
|
||||
class EngineTests(unittest.TestCase):
|
||||
@classmethod
|
||||
def setUpClass(cls):
|
||||
cls.network = patch('socket.socket.connect', side_effect=AssertionError('network forbidden'))
|
||||
cls.network.start()
|
||||
cls.addClassCleanup(cls.network.stop)
|
||||
if Engine:
|
||||
cls.engine = Engine(RESOURCES)
|
||||
|
||||
def setUp(self):
|
||||
self.assertIsNotNone(Engine, 'English engine has not been implemented')
|
||||
|
||||
def test_unicode_partition_and_three_offsets(self):
|
||||
for text in ['', ' \t\r\n', " She went!\r\nDogs’ paws\tcan't. e\u0301 café 😀 中文\u00a0\u200bend ",
|
||||
'well-known mother-in-law 👩💻 👨👩👧👦 🏳️🌈']:
|
||||
with self.subTest(text=text):
|
||||
result = self.engine.analyze(text)
|
||||
self.assertEqual(result['status'], 'ok')
|
||||
self.assertEqual(result['contract_version'], 'english-spike-v1')
|
||||
self.assertEqual(result['original_text'], text)
|
||||
self.assertEqual(result['text_sha256'], hashlib.sha256(text.encode()).hexdigest())
|
||||
tokens = result['tokens']
|
||||
self.assertEqual(''.join(t['text'] for t in tokens), text)
|
||||
cursor = 0
|
||||
for token in tokens:
|
||||
self.assertEqual(token['start_cp'], cursor)
|
||||
cursor = token['end_cp']
|
||||
self.assertGreater(cursor, token['start_cp'])
|
||||
self.assertEqual(text[token['start_cp']:cursor], token['text'])
|
||||
for encoding, unit, suffix in [('utf-8', 1, 'utf8'), ('utf-16-le', 2, 'utf16')]:
|
||||
start, end = token['start_' + suffix], token['end_' + suffix]
|
||||
self.assertEqual(text.encode(encoding)[start*unit:end*unit].decode(encoding), token['text'])
|
||||
self.assertEqual(len(text[:token['start_cp']].encode(encoding)) // unit, start)
|
||||
self.assertIn(token['kind'], ['word', 'space', 'punctuation'])
|
||||
self.assertEqual(cursor, len(text))
|
||||
|
||||
def test_contextual_irregular_lemma(self):
|
||||
tokens = self.engine.analyze('She went home. The children ate apples.')['tokens']
|
||||
lemmas = {t['text']: t['lemma'] for t in tokens}
|
||||
self.assertEqual(lemmas['went'], 'go')
|
||||
self.assertEqual(lemmas['children'], 'child')
|
||||
self.assertEqual(lemmas['ate'], 'eat')
|
||||
|
||||
def test_exact_then_explicit_lemma(self):
|
||||
exact = self.engine.lookup('DOG')
|
||||
self.assertEqual(exact['status'], 'exact')
|
||||
self.assertEqual(exact['matched_form'], 'dog')
|
||||
self.assertTrue(exact['entries'])
|
||||
self.assertLessEqual(len(exact['entries']), 12)
|
||||
self.assertEqual(exact, self.engine.lookup('DOG'))
|
||||
self.assertEqual(self.engine.lookup('went')['status'], 'not_found')
|
||||
lemma = self.engine.lookup('went', 'go')
|
||||
self.assertEqual(lemma['status'], 'lemma')
|
||||
self.assertEqual(lemma['matched_form'], 'go')
|
||||
self.assertEqual(self.engine.lookup('zzzxqvfiction')['status'], 'not_found')
|
||||
|
||||
def test_validation(self):
|
||||
for text in ['x' * 100001, '\ud800']:
|
||||
with self.assertRaises(ValueError):
|
||||
self.engine.analyze(text)
|
||||
with self.assertRaises(TypeError):
|
||||
self.engine.analyze(None)
|
||||
with self.assertRaises(ValueError):
|
||||
self.engine.lookup('\udfff')
|
||||
|
||||
def test_missing_resources_are_not_misses(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
engine = Engine(Path(directory), model_name='nonexistent_english_spike_model')
|
||||
self.assertEqual(engine.lookup('dog')['status'], 'resource_missing')
|
||||
self.assertEqual(engine.lookup('dog')['resource'], 'wordnet')
|
||||
with self.assertRaises(ResourceMissing) as error:
|
||||
engine.analyze('dog')
|
||||
self.assertEqual(error.exception.resource, 'model')
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
unittest.main()
|
||||
@@ -0,0 +1,14 @@
|
||||
export function validateTokens(text, tokens) {
|
||||
let end=0
|
||||
for (const token of tokens) {
|
||||
if (!Number.isInteger(token.start_utf16) || !Number.isInteger(token.end_utf16) || token.start_utf16!==end || token.end_utf16<=end || text.slice(token.start_utf16,token.end_utf16)!==token.text) throw new Error('原文位置校验失败')
|
||||
end=token.end_utf16
|
||||
}
|
||||
if(end!==text.length) throw new Error('原文还原失败')
|
||||
return true
|
||||
}
|
||||
|
||||
export function latestOnly() {
|
||||
let version=0
|
||||
return {next:()=>++version,current:value=>value===version}
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
import test from 'node:test'
|
||||
import assert from 'node:assert/strict'
|
||||
import { validateTokens, latestOnly } from './view.mjs'
|
||||
|
||||
test('UTF-16 positions reconstruct emoji and combining characters without normalization', () => {
|
||||
const text = '🙂 Café'
|
||||
const tokens = [{text:'🙂',start_utf16:0,end_utf16:2},{text:' ',start_utf16:2,end_utf16:3},{text:'Café',start_utf16:3,end_utf16:8}]
|
||||
assert.equal(validateTokens(text,tokens), true)
|
||||
assert.throws(() => validateTokens(text,[{text:'🙂',start_utf16:0,end_utf16:1}]))
|
||||
assert.throws(() => validateTokens(text,[]))
|
||||
})
|
||||
test('old analysis and lookup responses cannot replace newer text or selection', () => {
|
||||
const gate=latestOnly()
|
||||
const first=gate.next(), second=gate.next()
|
||||
assert.equal(gate.current(first),false)
|
||||
assert.equal(gate.current(second),true)
|
||||
gate.next()
|
||||
assert.equal(gate.current(second),false)
|
||||
})
|
||||
Reference in New Issue
Block a user